LATIDIA · Ciberseguridad
Reactivación de la alineación: defensa de los LLM de las fugas de cárcel a través de la coincidencia de entrada-salida consciente de la intención
arXiv: 2610.04470v1Announce Type: cross Resumen: Los modelos de lenguaje grandes (LLM) siguen siendo vulnerables a los ataques de jailbreak que ocultan intenciones dañinas dentro de mensajes contradictorios complejos. Las defensas existentes se basan principalmente en
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.04470v1 Announce Type: cross Abstract: Large language models (LLMs) remain vulnerable to jailbreak attacks that conceal harmful intent within complex adversarial prompts. Existing defenses primarily rely on input perturbation or harmful-output suppression, but they rarely model where malicious intent resides, resulting in brittle protection and excessive over-refusal. We propose SENTINEL, a plug-and-play, generation-time jailbreak defense that reframes mitigation as an intent extraction problem. Our key insight is that instruction-tuned LLMs exhibit strong input--output semantic consistency: regardless of jailbreak complexity, generated outputs tend to align with the attacker's true intent. SENTINEL exploits this property by matching semantically aligned input--output regions to extract intention-revealing subsequences, scores these subsequences using refusal-direction projections to