LATIDIA · Ciberseguridad
Pregúntele al experto: Aprendizaje de refuerzo guiado por LLM para defensa cibernética autónoma
arXiv: 2610.09337v1Announce Type: new Abstract: Los enfoques de aprendizaje por refuerzo (RL) basados en políticas han producido resultados prometedores para la defensa cibernética autónoma; sin embargo, son ineficientes para la muestra en entornos donde def
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.09337v1 Announce Type: new Abstract: Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semantically about security state space, high latency and trust assumptions prevent attractive in-line deployment models. We introduce Ask the Expert, a training-time guidance framework which first summarizes hard cyber-defense states, then intermittently queries an LLM for host-level defensive recommendations via a constrained action interface, and finally transforms those recommendations into tiered reward shaping for use with PPO. Because the LLM is discarded after training, deployment