LATIDIA · Ciberseguridad
Defensas universales para agentes de LLM integrados en herramientas contra ataques adversarios
arXiv: 2609.16098v1Tipo de anuncio: nuevo Resumen: los agentes del Modelo de Lenguaje Grande (LLM) han demostrado capacidades impresionantes en una variedad de dominios, particularmente cuando se integran con herramientas externas para tareas de varios pasos
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.16098v1 Announce Type: new Abstract: Large Language Model (LLM) agents have demonstrated impressive capabilities across a variety of domains, particularly when integrated with external tools for multi-step task completion. However, they are increasingly vulnerable to adversarial attacks, including direct prompt injection, indirect prompt injection, memory poisoning, and backdoor attacks, which exploit the model's openness to prompt injection and tool manipulation. In this work, we explore practical and generalizable defense strategies within a unified framework across these four attack types. We introduce two universal tool-based defenses: Attacker Tool Filtering, which uses anomaly detection (e.g., Isolation Forest) to identify and remove suspicious tools, and Normal Tool Recalling, a white-box method that restores