LATIDIA · Investigación
StoreBench: un entorno de comercio en vivo para evaluar y capacitar a agentes operadores autónomos
arXiv: 2610.10942v1Tipo de anuncio: nuevo Resumen: Los entornos de aprendizaje por refuerzo son ahora una palanca principal para mejorar las capacidades del modelo de lenguaje grande (LLM) en la post-entrenamiento, sin embargo, la mayoría de los puntos de referencia agénticos siguen siendo estadísticos
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.10942v1 Announce Type: new Abstract: Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty. Customers order around the clock, suppliers reprice and fail, and market shocks arrive with partial or no warning. The agent acts through the same 29 merchant tools a human operator would use,