LATIDIA · Investigación
El Atlas de Pareto de Ingeniería de Inferencia: ¿Qué optimizaciones dominan la frontera del costo, la calidad y la latencia?
arXiv:2609.17863v1 Announce Type: new Resumen: Las optimizaciones de inferencia de LLM informan sobre aceleraciones en diferentes modelos, GPU, indicaciones y métricas de calidad, lo que dificulta su comparación o combinación. Construimos un coste, calidad y
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.17863v1 Announce Type: new Abstract: LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine. We build a cost, quality, and latency Pareto atlas to identify the best configurations for different deployment constraints. Since exhaustive testing is impractical, we measure 54 configurations of Qwen2.5-7B-Instruct running on vLLM 0.12 across L4, A100, and H100 GPUs and use these anchors to calibrate a simulator. It reproduces measurements at anchored batch sizes, with cross campaign drift below 1.5 percent. A separate quality evaluation tests FP16, AWQ 4bit, FP8 weights, and FP8 KV cache on 200 GSM8K questions with five examples per prompt. Sparse