LATIDIA · Ciberseguridad
La seguridad no se compone: estado de bucle no decaído para agentes autónomos de LLM
arXiv: 2608.27141v5Announce Type: replace Resumen: Los agentes de modelos de lenguaje grandes se implementan cada vez más como bucles autónomos. Partiendo de un objetivo humano, tal sistema descubre repetidamente el trabajo, planifica, ejecuta la herramienta c
WhatsApp ↗Telegram ↗
La noticia
arXiv:2608.27141v5 Announce Type: replace Abstract: Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result is a separation: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because