UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Ciberseguridad

RADNPO: Optimización de preferencias negativas adaptativas sin referencias para el desaprendizaje de LLM

arXiv: 2609.34251v1Tipo de anuncio: nuevo Resumen: los modelos de lenguaje grandes (LLM) pueden memorizar contenido sensible, privado o con derechos de autor durante la capacitación previa, lo que hace que el desaprendizaje automático sea necesario para eliminar el conocimiento específico

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con RADNPO: Optimización de preferencias negativas adaptativas sin referencias para el desaprendizaje de LLM
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2609.34251v1 Announce Type: new Abstract: Large language models (LLMs) can memorize sensitive, private, or copyrighted content during pre-training, making machine unlearning necessary for removing targeted knowledge. Recent preference optimization (PO)-based unlearning methods improve stability over gradient ascent (GA)-based methods by introducing alignment-style objectives, which effectively suppress the probability of forget targets. However, target suppression alone does not sufficiently constrain the next-token distribution after unlearning. Existing methods provide limited control over how suppressed probability mass is redistributed and insufficiently adapt forgetting strength to target confidence and distributional concentration. Even after target suppression, probability mass may remain concentrated on a few non-target tokens, potentially producing repetitive or uninformative outputs. To address these

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto