LATIDIA · Ciberseguridad
CorrectGuard: Eyes-Off Correctness Estimation for Black-Box Security Guardrails
arXiv: 2610.03470v1 Tipo de anuncio: nuevo Resumen: los servicios de IA dependen cada vez más de las barandillas de seguridad de la caja negra, sin embargo, los regímenes de auditoría de modelos que preservan la privacidad a menudo no pueden medir qué tan bien funcionan estos sistemas tanto en un
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.03470v1 Announce Type: new Abstract: AI services increasingly rely on black-box security guardrails, yet privacy-preserving model auditing regimes often cannot measure how well these systems perform in both a human eyes-off production setting, which disallows human inspection of user input, and a machine eyes-off setting, which disallows model inspection of such input. We introduce CorrectGuard, an eyes-off correctness estimation framework for both settings, which involves an independent model-based evaluator predicting whether guardrail decisions on human- and machine-inaccessible inputs are correct using only labeled eyes-on data and without access to the guardrail's internals. We evaluate in-context learning, embedding, and finetuning-based correctness models under leave-one-dataset-out evaluation across 13 safety and security datasets