LATIDIA · Investigación
RideWay: Evaluación comparativa de la finalización eficiente de tareas para agentes lingüísticos que utilizan herramientas
arXiv:2609.17985v1 Announce Type: new Resumen: Los agentes de IA generalmente se evalúan por si completan una tarea. En la configuración de servicios interactivos, un agente exitoso aún puede frustrar a los usuarios haciendo preguntas repetidas,
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant searches, or making avoidable revisions. We introduce RideWay, an efficiency-centered benchmark for ridehailing agents in a stateful tool-calling environment, together with Efficiency Utility, a success-gated metric that discounts successful trajectories for excess tool calls and user-facing turns relative to task-specific reference effort. Human paired preferences calibrate the relative penalties, reflecting an aggregate service-workflow trade-off: extra dialogue often creates visible friction, whereas extra tool use can sometimes verify constraints or preserve user intent. Across 58 tasks and 24