LATIDIA · Robótica
GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives
arXiv: 2609.39601v1Tipo de anuncio: Cross Resumen: La conexión a tierra precisa es importante. Especifica qué objeto es el objetivo y dónde está ese objeto, incluso en desorden y para objetos pequeños, y tiene que ser lo suficientemente rápido como para cerrar
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.39601v1 Announce Type: cross Abstract: Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual