LATIDIA · Robótica
SynthDemo-RL: Rompiendo la barrera de recompensa cero en la adaptación de VLA con demostraciones sintéticas guiadas por LLM
arXiv: 2609.21650v1Announce Type: new Abstract: Fine-tuning Vision-Language-Action (VLA) models comúnmente se basa en demostraciones de teleoperación humana, mientras que el aprendizaje por refuerzo (RL) con escasas recompensas binarias se enfrenta a un
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.21650v1 Announce Type: new Abstract: Fine-tuning Vision-Language-Action (VLA) models commonly relies on human teleoperation demonstrations, while reinforcement learning (RL) with sparse binary rewards faces an exploration challenge when successful trajectories are rarely sampled. We propose SynthDemo-RL, a teacher-student framework in which an automated teacher converts simulator-privileged state into successful manipulation trajectories, a VLA student is distilled from them by supervised fine-tuning (SFT), and PPO with binary task-success rewards refines the student. We study reward coverage, the fraction of tasks for which at least one success is observed under the fixed evaluation protocol, as a complement to the average success rate. On LIBERO-PRO, a public benchmark of perturbed LIBERO tasks for