LATIDIA · Ciberseguridad
KUDA: Desaprendizaje de conocimientos mediante la representación desviada para modelos de lenguaje grandes
arXiv: 2602.19275v3Tipo de anuncio: reemplazar Resumen: Los modelos de lenguaje grandes (LLM) adquieren una gran cantidad de conocimiento a través de la capacitación previa en cuerpos vastos y diversos. Si bien esto dota a los LLM de sólidas capacidades en ge
WhatsApp ↗Telegram ↗
La noticia
arXiv:2602.19275v3 Announce Type: replace Abstract: Large language models (LLMs) acquire a large amount of knowledge through pre-training on vast and diverse corpora. While this endows LLMs with strong capabilities in generation and reasoning, it amplifies risks associated with sensitive, copyrighted, or harmful content in training data. LLM unlearning, which aims to remove specific knowledge encoded within models, is a promising technique to reduce these risks. However, existing LLM unlearning methods often force LLMs to generate random or incoherent answers due to their inability to alter the encoded knowledge precisely. To achieve effective unlearning at the knowledge level of LLMs, we propose Knowledge Unlearning by Deviating representAtion (KUDA). We first utilize