LATIDIA · Ciberseguridad
Aprobación de la prueba en la que se entrenó: reevaluación de los detectores de inyección rápida para agentes de LLM
arXiv: 2610.03448v1Tipo de anuncio: nuevo Resumen: los agentes de LLM examinan cada vez más los resultados de las herramientas con pequeños detectores de inyección rápida, y los equipos eligen entre los detectores por sus puntajes en los puntos de referencia públicos. Preguntamos si tho
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.03448v1 Announce Type: new Abstract: LLM agents increasingly screen tool outputs with small prompt-injection detectors, and teams choose among detectors by their scores on public benchmarks. We ask whether those scores predict how a detector behaves inside an agent. We replay the ground-truth tool calls of two agent benchmarks, AgentDojo and tau-bench, without an LLM to obtain tool outputs that are benign by construction, label injected outputs by differential replay, and evaluate fifteen detectors, including Meta's Prompt Guard 2, and two task-aware LLM judges on these outputs and on the BIPIA benchmark. Detection rankings transfer poorly between benchmarks: the best detector on BIPIA catches 2% of AgentDojo injections at a