UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Investigación

Razonamiento-Preservación de ajuste fino de LLM posteriores a RL con LoRA de base nula

arXiv:2609.25618v1 Announce Type: new Abstract: Reinforcement learning (RL)-based post-training has become an effective approach for eliciting reasoning capabilities in large language models (LLMs). Sin embargo, la adaptación de la publicación

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con Razonamiento-Preservación de ajuste fino de LLM posteriores a RL con LoRA de base nula
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2609.25618v1 Announce Type: new Abstract: Reinforcement learning (RL)-based post-training has become an effective approach for eliciting reasoning capabilities in large language models (LLMs). However, adapting post-RL models to new knowledge domains or behaviors through subsequent supervised fine-tuning (SFT) can severely overwrite these capabilities. Existing approaches mitigate such forgetting through experience replay, specialized initialization, or constrained optimization using gradient projection, but either provide limited preservation or incur substantial training overhead. Our analysis shows that reasoning activations concentrate in low-dimensional subspaces, leaving substantial null-space capacity for adaptation, and that the corresponding approximate null spaces can be reliably estimated from a modest number of examples. Motivated by these observations, we propose Null-Basis Low-Rank

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto