LATIDIA · Ciberseguridad
CoDeL: Defensa coevolucionaria contra la inyección rápida indirecta en agentes basados en LLM
arXiv: 2609.34463v1Tipo de anuncio: nuevo Resumen: los agentes basados en el modelo de lenguaje grande (LLM) dependen cada vez más de herramientas y contenido externos, exponiéndolos a una inyección rápida indirecta (IPI). Esta amenaza ha motivado una amplia
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.34463v1 Announce Type: new Abstract: Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based defenses are typically optimized on a static distribution of explicit injections. They learn surface-form cues rather than the boundary between serving the user and obeying an injected objective, and therefore fail when malicious intent is folded into a plausible workflow and deferred for several turns. We present CoDeL, a defense that hardens agent against an attack distribution it reshapes as it trains. The defender