LATIDIA · Ciberseguridad
Protección de los LLM a través de señales de seguridad latentes de diagnóstico de modelo de Dark Knowledge
arXiv:2610.07532v1 Tipo de anuncio: nuevo Resumen: los LLM han avanzado rápidamente, lo que aumenta las preocupaciones sobre su seguridad. Trabajos recientes han propuesto enfoques para detectar y defenderse contra ataques, incluidas las defensas en deco
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.07532v1 Announce Type: new Abstract: LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and