LATIDIA · Ciberseguridad
Modelado probabilístico de jailbreak en LLMs multimodales: de la cuantificación a la aplicación
arXiv:2503.06989v5 Announce Type: replace Resumen: Recientemente, los Modelos Multimodales de Lenguaje Grande (MLLM) han demostrado su capacidad superior para comprender el contenido multimodal. Sin embargo, siguen siendo vulnerables a la cárcel
WhatsApp ↗Telegram ↗
La noticia
arXiv:2503.06989v5 Announce Type: replace Abstract: Recently, Multimodal Large Language Models (MLLMs) have demonstrated their superior ability in understanding multimodal content. However, they remain vulnerable to jailbreak attacks, which exploit weaknesses in their safety alignment to generate harmful responses. Previous studies categorize jailbreaks as successful or failed based on whether responses contain malicious content. However, given the stochastic nature of MLLM responses, this binary classification of an input's ability to jailbreak MLLMs is inappropriate. Derived from this viewpoint, we introduce jailbreak probability to quantify the jailbreak potential of an input, which represents the likelihood that MLLMs generated a malicious response when prompted with this input. We approximate this probability through multiple