LATIDIA · Robótica
Mejorar la semejanza humana en los agentes de aprendizaje de refuerzo a través de la cuantificación jerárquica de macroacciones
arXiv: 2605.30928v2Announce Type: replace Resumen: Los agentes similares a los humanos son un objetivo de larga data de la inteligencia artificial. A pesar de un buen desempeño, la mayoría de los agentes de aprendizaje por refuerzo (RL) siguen siendo impulsados por las recompensas y a menudo
WhatsApp ↗Telegram ↗
La noticia
arXiv:2605.30928v2 Announce Type: replace Abstract: Human-like agents are a long-standing goal of artificial intelligence. Despite strong performance, most reinforcement learning (RL) agents remain reward-driven and often exhibit behaviors that differ from humans, limiting interpretability and reliability. In this work, we introduce a novel human-like RL framework that predicts action sequences closely aligned with human behaviors while maximizing rewards. Specifically, we encode human demonstrations into macro actions using a hierarchical macro action quantization approach (HiMAQ) consisting of two successive levels of vector quantization. The lower quantization level maps input actions to fine-grained subaction clusters, while the higher quantization level aggregates these subaction clusters into action clusters. Extensive evaluations on the D4RL