LATIDIA · Robótica
Adaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering
arXiv: 2609.21502v1Announce Type: cross Resumen: Los modelos de cimentación 3D recientes permiten el razonamiento geométrico generalizable a partir de imágenes RGB, pero siguen siendo limitados en memoria persistente, escalabilidad y modelado de escenas renderizables.
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.21502v1 Announce Type: cross Abstract: Recent 3D foundation models enable generalizable geometric reasoning from RGB images but remain limited in persistent memory, scalability, and renderable scene modeling. We present a memory-centric 3D foundation model for scalable robotic localization, reconstruction, and Gaussian rendering. Its core is an adaptive world memory mechanism that combines transformer-based gated updates with test-time temporal-spatial regulation. Learned gates control recurrent memory propagation, while temporal state evolution and spatial observation-state consistency regulate token-wise updates and forgetting over long image sequences. To support large-scale mapping, we organize memory into local submaps and integrate progressive mapping and tracking, loop closure, and SL(4)-based global refinement to maintain local accuracy and global