LATIDIA · Ciberseguridad
Defensa Autónoma: Aprendizaje Continuo de Políticas de Seguridad para Agentes LLM
arXiv:2609.36603v1 Anuncio Tipo: nuevo Resumen: Los modelos de lenguaje grandes (LLM) impulsan cada vez más a los agentes que acceden a información confidencial, utilizan herramientas externas y modifican repositorios de software. Aunque estas capacidades
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.36603v1 Announce Type: new Abstract: Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and vulnerable code generation. Existing defenses often require retraining, fail to adapt to evolving attacks, or address only a single threat pattern. To address these limitations, we propose Self-Evolving Defense (SED), a training-free framework that distills harmful agent trajectories into reusable security policies without updating model weights. By retrieving relevant policies for future tasks, SED continually adapts to new attacks while retaining knowledge across attack scenarios. To evaluate the effectiveness of