UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Robótica

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

arXiv: 2610.09455v1Tipo de anuncio: Cross Resumen: Recientemente, los enfoques que aprovechan los conjuntos de datos de video humano para la capacitación en políticas de robots se han vuelto cada vez más frecuentes. Sin embargo, la mayoría de los rastreadores de mano existentes regresan a la pose fr

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2610.09455v1 Announce Type: cross Abstract: Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning,

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto