UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Ciberseguridad

Reflexiones y fragmentos: asegurar los LLM contra ataques secuenciales de mosaico

arXiv: 2610.05346v1Announce Type: cross Abstract: Self-play red-teaming mejora la seguridad del modelo de lenguaje al enfrentar los roles de atacante y defensor entre sí en un juego de suma cero. Sin embargo, los adversarios reales cada vez más

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con Reflexiones y fragmentos: asegurar los LLM contra ataques secuenciales de mosaico
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2610.05346v1 Announce Type: cross Abstract: Self-play red-teaming improves language-model safety by pitting attacker and defender roles against each other in a zero-sum game. However, real adversaries increasingly use mosaic attacks: multi-turn sequences whose individual fragments are innocuous in isolation yet assemble into a harmful payload. We develop a theory of mosaic defense that characterizes what is required to prevent such attacks without sacrificing helpfulness. We first show that no fixed bounded window of recent prompts is sufficient in general: safety-relevant information may occur arbitrarily far back in the interaction. We formalize a watchman, an online state mechanism that carries this information forward, and show that under explicit assumptions it enables

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto