LATIDIA · Investigación
Cuando la pérdida de reconstrucción más baja duele: refinamiento distribucionalmente robusto para la cuantificación de LLM de baja bits
arXiv: 2610.11226v1Announce Type: new Resumen: La cuantificación post-entrenamiento (PTQ) solo con peso se basa en gran medida en la minimización de la pérdida de reconstrucción para preservar la calidad del modelo con baja precisión. Mostramos que los pesos favorecidos
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.11226v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) relies heavily on reconstruction loss minimization to preserve model quality at low precision. We show that the weights favored by minimizing this loss need not yield better model performance on new tasks. In fact, we find that lower reconstruction loss can even degrade model performance on the same calibration data. Our analysis further shows that weights with lower reconstruction loss on calibration data can have higher loss than other weights when the distribution of input activations changes. Motivated by these observations and our analysis, we propose Distributionally Robust Quantization (DRQ), a post-hoc refinement process that minimizes worst-case reconstruction loss over a