UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Investigación

Fed-GRPO: Optimización de políticas relativas de grupos federados impulsada por señales de recompensa

arXiv: 2610.11502v1Tipo de anuncio: nuevo Resumen: los modelos de lenguaje grandes (LLM) han demostrado fuertes capacidades de razonamiento cuando se ajustan con aprendizaje de refuerzo (RL), particularmente a través de Group Relative Policy Optimizat

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con Fed-GRPO: Optimización de políticas relativas de grupos federados impulsada por señales de recompensa
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2610.11502v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL), particularly through Group Relative Policy Optimization (GRPO). However, existing GRPO methods assume centralized access to training data, which may not hold in practice due to privacy or regulatory constraints. To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication. Fed-GRPO contains three reward-signal-driven mechanisms: (i) \emph{signal-weighted aggregation} that weights clients by their reward standard deviation,

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto