LATIDIA · Ciberseguridad
Renderizar antes de leer: la representación visual como defensa contra la inyección de mensajes
arXiv:2609.36121v1 Anuncio Tipo: nuevo Resumen: Los modelos de lenguaje grandes son vulnerables a ataques de inyección de mensajes, donde contenido adversario de terceros puede secuestrar el comportamiento del modelo. En este artículo, estudiamos el papel de la inyección de mensajes.
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.36121v1 Announce Type: new Abstract: Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to obey textual instructions while treating other modalities mainly as content to parse or describe. We then demonstrate how this gap can be turned into