LATIDIA · Ciberseguridad
Speedbumps: Ataques de rechazo a la decodificación especulativa
arXiv: 2610.10929v1Tipo de anuncio: nuevo Resumen: La decodificación especulativa es una técnica popular para aumentar la velocidad y reducir los costos de la inferencia del modelo de lenguaje grande (LLM) mediante la verificación de múltiples tokens de borrador en un
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.10929v1 Announce Type: new Abstract: Speculative decoding is a popular technique for increasing the speed and reducing the costs of large language model (LLM) inference by verifying multiple draft tokens in a single target-model forward pass. The resulting benefit depends on the ability of the drafter to approximate the target model's distribution. In this work, we study Speculative Rejection Attacks (SRAs), a novel class of attacks that cause draft and target models to disagree more often, resulting in fewer draft tokens being accepted per draft cycle. This leads to more target model forward passes needed per generated token, slowing down inference and increasing costs for the victim. We introduce two