LATIDIA · Ciberseguridad
Detección válida en cualquier momento de la exfiltración de peso LLM
arXiv: 2610.11843v1Tipo de anuncio: nuevo Resumen: Un servidor de inferencia LLM comprometido puede filtrar los pesos del modelo codificando bits de carga útil en opciones de token que de otro modo serían plausibles. Una repetición del mismo mensaje en un servidor de confianza puede
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.11843v1 Announce Type: new Abstract: A compromised LLM inference server can leak model weights by encoding payload bits in otherwise plausible token choices. A replay of the same prompt in a trusted server can expose such deviations, but benign numerical nondeterminism also causes token mismatches. Patient attackers can therefore hide within normal variation unless evidence is combined across responses. We introduce a prompt-level e-process that calibrates whole-response mismatch events on trusted benign traffic and accumulates evidence sequentially while, under a calibration-transfer assumption, controlling the probability of any false alarm over an unbounded monitoring horizon. We evaluate it on four models against a seed-blind attack and a stronger seed-aware attack that