LATIDIA · Ciberseguridad
Reflexiones y fragmentos: asegurar los LLM contra ataques secuenciales de mosaico
arXiv: 2610.05346v1Announce Type: cross Abstract: Self-play red-teaming mejora la seguridad del modelo de lenguaje al enfrentar los roles de atacante y defensor entre sí en un juego de suma cero. Sin embargo, los adversarios reales cada vez más
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.05346v1 Announce Type: cross Abstract: Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isolation yet assemble into a harmful payload. We develop a theory of mosaic defense that characterizes what is required to prevent such attacks without sacrificing helpfulness. We first show that no fixed bounded window of recent prompts is sufficient in general: safety-relevant information may occur arbitrarily far back in the interaction. We formalize a watchman, an online state mechanism that carries this information forward, and show that under explicit assumptions it enables