LATIDIA · Ciberseguridad
Extracción de la AGUJA en el pajar: eliminación de la puerta trasera en LLM a través de la ortogonalización del peso
arXiv: 2610.00348v1Tipo de Anuncio: nuevo Resumen: Los ataques de puerta trasera se pueden implantar en Modelos de Lenguaje Grande (LLM) durante el entrenamiento, causando un comportamiento no deseado cuando aparece un disparador en la entrada. Defensa de puerta trasera existente
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.00348v1 Announce Type: new Abstract: Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original