LATIDIA · Investigación
Volver a la definición: Estimación de las ventajas de nivel escalonado a través de gráficos de trayectoria para el aprendizaje de refuerzo genético
arXiv: 2609.28963v1Tipo de anuncio: nuevo Resumen: Los métodos de aprendizaje por refuerzo (RL) basados en grupos, como GRPO y sus variantes, se han convertido en un paradigma líder para el razonamiento formativo y los modelos de lenguaje grande (LLM)
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.28963v1 Announce Type: new Abstract: Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since coarse-grained trajectory-level advantages are hard to accurately reflect the contribution of individual steps (i.e, failed trajectories may contain valuable steps). Revisiting the foundational RL definition, we notice that GRPO's success on single-turn tasks stems from its advantage estimation strategy, which adheres to the basic definition: the mean reward of multiple actions sampled from the same state constitutes a credible state-value