LATIDIA · Investigación
¿Qué esperamos de los LLM? Mapeo del diseño de puntos de referencia de LLM
arXiv:2609.19182v1 Tipo de anuncio: nuevo Resumen: Los puntos de referencia son fundamentales para evaluar y comunicar el progreso en los modelos de lenguaje grandes (LLM). Sin embargo, las clasificaciones de modelos por sí solas revelan poco sobre cómo el requisito de evaluación
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.19182v1 Announce Type: new Abstract: Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer