LATIDIA · Robótica
Aprendizaje en contexto de vídeo recursivo para robots genéticos
arXiv: 2610.06843v1Tipo de anuncio: nuevo Resumen: los agentes de LLM que organizan las políticas de visión-lenguaje-acción (VLA) congeladas mejoran en todos los episodios a través de la memoria de texto, que registra lo que hizo el agente pero no cómo la tarea
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.06843v1 Announce Type: new Abstract: LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built