LATIDIA · Robótica
PhysBrain 1.5: de los modelos de visión y lenguaje a los modelos de fundamentos físicos
arXiv:2609.14973v1 Tipo de anuncio: Cross Resumen: Presentamos PhysBrain 1.5, un modelo unificado para comprender entornos físicos, generar acciones y predecir estados futuros. Motivado por el bucle físico de obs
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.14973v1 Announce Type: cross Abstract: We present PhysBrain 1.5, a unified model for understanding physical environments, generating actions, and predicting future states. Motivated by the physical loop of observation, interaction, and environmental change, we bring these capabilities into a common learning framework. Starting from a general vision--language model, we encode language responses, end-effector motion, and dense visual targets as discrete sequences and jointly optimize them with autoregressive next-token prediction. Pre-training draws its embodied supervision entirely from human interaction videos, using task-centered episodes to pair semantic and spatial context with recovered motion and subsequent observations. We then adapt the model through supervised fine-tuning on a mixture of human demonstrations, robot trajectories,