UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Investigación

DEEPO: Optimización de políticas mejorada de doble entropía para la alucinación en MLLM

arXiv:2609.28570v1 Announce Type: new Resumen: El aprendizaje por refuerzo (RL) se utiliza ampliamente para agudizar el razonamiento en modelos multimodales de lenguaje grande (MLLM), pero su efecto sobre la alucinación es desigual. Rastreamos esto a dos

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con DEEPO: Optimización de políticas mejorada de doble entropía para la alucinación en MLLM
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2609.28570v1 Announce Type: new Abstract: Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{correction chain} from reward to parameter update. At the rollout level, hard queries---those with high semantic entropy---frequently produce unanimously wrong sample groups, collapsing the group-relative advantage to zero exactly where hallucination risk is highest. At the optimization level, confident-but-wrong tokens are gradient-invisible: a categorical policy's expected score-gradient norm vanishes as its distribution sharpens, so the predictions that most need correction receive the weakest updates. We propose Dual-Entropy Enhanced Policy Optimization (DEEPO), a dual-stage enhancement

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto