UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Investigación

Aprendizaje de preferencias heterogéneas

arXiv: 2609.17847v1Tipo de anuncio: nuevo Resumen: Aprender de la retroalimentación humana se ha convertido en un paradigma central para entrenar sistemas modernos de IA, donde los modelos de utilidad humana se utilizan como modelos de recompensa en el aprendizaje de políticas.

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con Aprendizaje de preferencias heterogéneas
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2609.17847v1 Announce Type: new Abstract: Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning. Existing methods typically assume a \emph{universal utility} function shared across a population and treat disagreement between annotators as stochastic variation. While suitable for objective tasks, this assumption breaks down in subjective domains where preferences vary systematically across individuals. We study the problem of subjective preference learning, in which observed choices arise from heterogeneous but internally consistent utility functions. Drawing upon rational choice theory, RCT \parencite{tversky1981framing}, we introduce \emph{individuated utility} functions conditioned on both the individual and their decision

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto