LATIDIA · Robótica
AR-WAM: un modelo de acción mundial preparado para agentes con acondicionamiento visual para la manipulación robótica
arXiv: 2609.23578v1Tipo de anuncio: nuevo Resumen: A medida que los agentes de IA se vuelven cada vez más capaces, el control robótico impulsado por agentes se está convirtiendo en un paradigma convincente. Sin embargo, los modelos predominantes de visión-lenguaje-acción (VLA) y WOR
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.23578v1 Announce Type: new Abstract: As AI agents become increasingly capable, agent-driven robotic control is emerging as a compelling paradigm. However, prevailing vision-language-action (VLA) models and world action models (WAMs) still rely on natural-language instructions to specify manipulation tasks, an ill-suited interface for agent-driven control: referentially ambiguous, spatially imprecise, redundant with the agent's inherent language understanding, and entangling intent with execution. We present AR-WAM, a visual-conditioned, agent-ready world action model that replaces language with two complementary conditions: a visual grounding prompt (a bounding box of the target) denoting the interaction object and location, and a learnable operation token dictating the atomic skill to execute. Our compact 0.5B-parameter model, with a