LATIDIA · Robótica
XEmbodied: un modelo de cimentación con señales geométricas y físicas mejoradas para entornos incorporados a gran escala
arXiv: 2604.18484v2Tipo de anuncio: reemplazar-cruz Resumen: Los modelos de Visión-Lenguaje-Acción (VLA) impulsan los sistemas autónomos de próxima generación, pero entrenarlos requiere anotaciones escalables y de alta calidad del entorno complejo
WhatsApp ↗Telegram ↗
La noticia
arXiv:2604.18484v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud pipelines rely on generic vision-language models (VLMs) that lack geometric reasoning and domain semantics due to their 2D image-text pretraining. To address this mismatch, we propose XEmbodied, a cloud-side foundation model that endows VLMs with intrinsic 3D geometric awareness and interaction with physical cues (e.g., occupancy grids, 3D boxes). Instead of treating geometry as auxiliary input, XEmbodied integrates geometric representations via a structured 3D Adapter and distills physical signals into context tokens using an Efficient Image-Embodied Adapter. Through progressive domain curriculum and reinforcement learning post-training, XEmbodied