LATIDIA · Ciberseguridad
Repensar la denegación de servicio de latencia: atacar el marco de servicio de LLM, no el modelo
arXiv:2602.07878v2 Announce Type: replace Resumen: La inferencia LLM es inherentemente costosa, incluso una desaceleración modesta puede traducirse en costos operativos sustanciales y graves riesgos de disponibilidad. Recientemente, un creciente cuerpo de
WhatsApp ↗Telegram ↗
La noticia
arXiv:2602.07878v2 Announce Type: replace Abstract: LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting inputs to trigger worst-case output lengths. However, we report a contrary finding that these algorithmic-level latency attacks are largely ineffective against modern LLM serving systems. We reveal that system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users. Thus, in this paper, we shift our focus from the algorithm to the system layer, and introduce a new Fill and Squeeze attack strategy targeting the state