LATIDIA · Ciberseguridad
¿Seguridad en los lotes? Comprender y mitigar las fallas de seguridad en las solicitudes de lotes
arXiv:2608.02681v2 Announce Type: replace Abstract: Batch prompting is a practical inference strategy for large language models, but its safety implications remain underexplored. Mostramos que el éxito de batch prompti
WhatsApp ↗Telegram ↗
La noticia
arXiv:2608.02681v2 Announce Type: replace Abstract: Batch prompting is a practical inference strategy for large language models, but its safety implications remain underexplored. We show that the success of batch prompting for utility does not extend to safety: a harmful question that is reliably refused in isolation can elicit a harmful response when embedded in a batch of benign questions. We identify this as a distinct safety failure mode -- not reducible to known vulnerabilities such as in-context learning or long-context effects -- and analyze its causes from two complementary perspectives: alignment signal weakening and refusal signal dilution. Across widely used open-source and frontier commercial models, batch prompting consistently achieves high