LATIDIA · Ciberseguridad
La frecuencia no es sensibilidad Identificación de expertos sensibles a la seguridad en Moe disperso LLM
arXiv: 2610.02910v1Tipo de anuncio: Cross Resumen: La supresión de un pequeño conjunto de expertos enrutados puede debilitar el comportamiento de seguridad de un modelo de lenguaje de mezcla de expertos (Moe) disperso sin reentrenamiento. Qué expertos suprimir
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.02910v1 Announce Type: cross Abstract: Suppressing a small set of routed experts can weaken the safety behavior of a sparse Mixture-of-Experts (MoE) language model without retraining. Which experts to suppress is therefore a security question, and the usual answer is activation frequency, but frequency measures use, not influence. We test an alternative: router-gradient sensitivity, the sensitivity of the sequence loss to the gate weights that select an expert. Across five MoE architectures, we rank experts by each signal on 500 benign and 500 malicious prompts and measure refusal on 100 held-out malicious prompts under two budgets: equal expert counts and equal nominal malicious routing traffic (1%-5%). Under each of the