LATIDIA · Robótica
RAÍZ: Descubrir recompensas por comportamientos incorporados especificados por el usuario
arXiv: 2610.04250v1Tipo de anuncio: nuevo Resumen: El aprendizaje de refuerzo para el control incorporado sigue limitado por la dificultad de la especificación de la recompensa. Aunque los métodos recientes basados en el modelo de lenguaje grande (LLM) pueden sintetizar
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.04250v1 Announce Type: new Abstract: Reinforcement learning for embodied control remains constrained by the difficulty of reward specification. Although recent large language model (LLM)-based methods can synthesize reward functions from natural-language descriptions, they often fail to capture subtle behavioral properties that humans care about, such as natural gait, posture, and movement style. This limitation arises because many desired behaviors are easier to recognize visually than to encode in a reward function. We introduce Reward Optimization via Observable Trees (ROOT), a framework for discovering reward functions that align learned policies with user-specified embodied behaviors. Rather than relying solely on scalar training statistics, ROOT casts reward design as an observation-guided search over