UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Robótica

SynthDemo-RL: Rompiendo la barrera de recompensa cero en la adaptación de VLA con demostraciones sintéticas guiadas por LLM

arXiv: 2609.21650v1Announce Type: new Abstract: Fine-tuning Vision-Language-Action (VLA) models comúnmente se basa en demostraciones de teleoperación humana, mientras que el aprendizaje por refuerzo (RL) con escasas recompensas binarias se enfrenta a un

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con SynthDemo-RL: Rompiendo la barrera de recompensa cero en la adaptación de VLA con demostraciones sintéticas guiadas por LLM
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2609.21650v1 Announce Type: new Abstract: Fine-tuning Vision-Language-Action (VLA) models commonly relies on human teleoperation demonstrations, while reinforcement learning (RL) with sparse binary rewards faces an exploration challenge when successful trajectories are rarely sampled. We propose SynthDemo-RL, a teacher-student framework in which an automated teacher converts simulator-privileged state into successful manipulation trajectories, a VLA student is distilled from them by supervised fine-tuning (SFT), and PPO with binary task-success rewards refines the student. We study reward coverage, the fraction of tasks for which at least one success is observed under the fixed evaluation protocol, as a complement to the average success rate. On LIBERO-PRO, a public benchmark of perturbed LIBERO tasks for

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto