LATIDIA · Ciberseguridad
Optimización de la dirección del señalizador: una defensa post-hoc contra la destrucción de LLM
arXiv: 2609.16204v1Tipo de Anuncio: Cross Resumen: Las barandillas de seguridad en modelos de lenguaje de peso abierto se pueden omitir fácilmente utilizando la Ablación de Funciones de Rechazo (RFA), una técnica que identifica y proyecta un rechazo lineal
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.16204v1 Announce Type: cross Abstract: Safety guardrails in open-weight language models can be readily bypassed using Refusal Feature Ablation (RFA), a technique that identifies and projects out a linear refusal direction from the residual stream, often achieving a high attack success rate (ASR) while preserving model capability. Defending against these attacks typically requires computationally expensive safety finetuning for every new checkpoint. We introduce Decoy Direction Optimization (DDO), a fast, post-hoc weight-editing defense that requires no base-model finetuning. Our approach is based on a simple mechanistic insight: ablation attacks rely on contrastive estimators to find the refusal direction. Rather than trying to hide the true refusal circuitry, DDO actively injects a