LATIDIA · Ciberseguridad
LEVANTADO: Autodestilación por robustez para impulsar la inyección en agentes de LLM
arXiv: 2610.06401v1Announce Type: new Resumen: Los agentes de modelos de lenguaje que utilizan herramientas son vulnerables a la inyección de mensajes indirectos porque deben actuar sobre contenido externo no confiable. Las defensas de tiempo de entrenamiento existentes pueden reducir
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.06401v1 Announce Type: new Abstract: Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defenses can reduce attack success rates, but often at the cost of general capabilities. We show that training-based defenses induce substantial drift in the model's output distribution, altering its behavior even in benign settings and providing a potential mechanism for utility degradation. We further identify a failure mode of these defenses: On benign tool-use tasks, the model refrains from a step needed to finish an authorized task, particularly when that step is indicated by a tool output. To address these limitations, we introduce RAISED (Robust Attack