LATIDIA · Ciberseguridad
Pisos falsos: las evaluaciones de enrutamiento de seguridad de LLM se interrumpen en el turno de distribución
arXiv:2610.01535v1 Tipo de anuncio: nuevo Resumen: Los routers de seguridad envían cada solicitud a uno de varios modelos y se juzgan contra el mejor modelo único. Un punto de referencia de enrutamiento importante elige ese comparador en la evaluación da
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.01535v1 Announce Type: new Abstract: Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered