LATIDIA · Ciberseguridad
Detección de inyección rápida para agentes de correo electrónico a través del modelado de la cadena de ataque
arXiv: 2609.30657v1Tipo de anuncio: nuevo Resumen: los asistentes de correo electrónico de modelo de lenguaje grande son particularmente vulnerables a la inyección de aviso indirecto porque el contenido de correo electrónico no confiable se puede recuperar en el contexto del modelo e i
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.30657v1 Announce Type: new Abstract: Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary malicious text classification, which overlooks the important factor that harmful agent behavior often arises through a sequence of stages. We propose a detection framework that models this attack chain by combining a text detector, verifiers specific to each stage, explicit rule-based risk signals, user intent and action consistency analysis, and a logistic decision policy. To support this framework, we derive attack chain labels from prompt injection datasets,