UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Investigación

RideWay: Evaluación comparativa de la finalización eficiente de tareas para agentes lingüísticos que utilizan herramientas

arXiv:2609.17985v1 Announce Type: new Resumen: Los agentes de IA generalmente se evalúan por si completan una tarea. En la configuración de servicios interactivos, un agente exitoso aún puede frustrar a los usuarios haciendo preguntas repetidas,

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con RideWay: Evaluación comparativa de la finalización eficiente de tareas para agentes lingüísticos que utilizan herramientas
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant searches, or making avoidable revisions. We introduce RideWay, an efficiency-centered benchmark for ridehailing agents in a stateful tool-calling environment, together with Efficiency Utility, a success-gated metric that discounts successful trajectories for excess tool calls and user-facing turns relative to task-specific reference effort. Human paired preferences calibrate the relative penalties, reflecting an aggregate service-workflow trade-off: extra dialogue often creates visible friction, whereas extra tool use can sometimes verify constraints or preserve user intent. Across 58 tasks and 24

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto