LATIDIA · Robótica
I-Perceive: un modelo básico para la percepción activa de la visión y el lenguaje
arXiv: 2603.00600v3Tipo de anuncio: reemplazar Resumen: La percepción activa, la capacidad de un robot para seleccionar proactivamente puntos de vista para adquirir información relevante para la tarea, es esencial para un funcionamiento sólido en el entorno del mundo real
WhatsApp ↗Telegram ↗
La noticia
arXiv:2603.00600v3 Announce Type: replace Abstract: Active perception - the ability of a robot to proactively select viewpoints to acquire task-relevant information - is essential for robust operation in real-world environments. However, existing approaches are typically limited to fixed objectives or constrained settings, and struggle to generalize to open-ended perception intents specified in natural language. We propose I-Perceive, a foundation model for language-conditioned active perception in large-scale indoor environments. Given a query image, a set of context images, and a natural language instruction, I-Perceive predicts a 6D camera pose that fulfills the specified perception intent. The model integrates a vision-language pathway for semantic grounding with a geometric reasoning pathway for multi-view