LATIDIA · Investigación
UniCAR-RL: Ver mejor antes de pensar más profundamente en matemáticas visuales
arXiv: 2609.13849v1Tipo de Anuncio: nuevo Resumen: Los Modelos Multimodales de Lenguaje Grande (MLLM) a menudo tienen dificultades con el razonamiento visual matemático complejo principalmente debido a la falta de percepción de grano fino, lo que causa una visión inicial
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.13849v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) often struggle with complex mathematical visual reasoning primarily due to a lack of fine-grained perception, causing initial visual hallucinations to directly trigger cascading reasoning failures. In traditional end-to-end reinforcement learning (RL), sparse rewards fail to decouple perceptual hallucinations from logical missteps, hindering targeted perception optimization. Alternatively, fine-tuning with perception-enhanced CoT data incurs high costs and hallucinations. In this paper, we address these challenges by proposing UniCAR-RL, an annotation-free RL framework. By explicitly decoupling the optimization of perception and reasoning during the training process, it achieves isolation and optimization of both capabilities. Specifically, UniCAR-RL consists of three synergistic branches: 1) a