LATIDIA · Ciberseguridad
ToolFence: Autorización de grano fino para agentes LLM que utilizan herramientas seguras
arXiv: 2609.37196v1Announce Type: new Resumen: Los agentes LLM que utilizan herramientas siguen siendo vulnerables a la inyección indirecta de mensajes porque las instrucciones confiables y los comentarios no confiables comparten un contexto, lo que permite que el contenido malicioso t
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.37196v1 Announce Type: new Abstract: Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical deployment. We introduce ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to