UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Ciberseguridad

Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning

arXiv: 2610.10655v1Tipo de anuncio: Cross Resumen: Los modelos de lenguaje grandes (LLM) inevitablemente internalizan cantidades sustanciales de información confidencial o privada durante la capacitación previa, mientras que el desaprendizaje de LLM tiene como objetivo selectivamente

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2610.10655v1 Announce Type: cross Abstract: Large Language Models (LLMs) inevitably internalize substantial amounts of sensitive or private information during pre-training, while LLM unlearning aims to selectively erase specific knowledge to prevent privacy leakage with minimal loss of model utility. However, existing methods struggle to balance forget quality with utility, and typically incur substantial computational costs due to parameter fine-tuning. To address this, we propose Nullify, a training-free, non-destructive activation steering method for LLM unlearning. Nullify employs steering vectors during inference to redirect privacy-related activations away from their memorized answers, while satisfying a null-space constraint that leaves retained-query activations essentially unaffected to maintain utility. Evaluations on TOFU and MUSE show that

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto