LATIDIA · Ciberseguridad
Modelos de decisión calibrados para arneses de prueba de penetración autónomos: JEV y Laya como capas de decisión del sistema uno para agentes Pentest impulsados por LLM
arXiv: 2609.28940v1Tipo de anuncio: nuevo Resumen: Los arneses autónomos de prueba de penetración utilizan modelos de lenguaje grande (LLM) para el reconocimiento, la explotación y la presentación de informes, pero a menudo se basan en esos mismos modelos para confirmar
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.28940v1 Announce Type: new Abstract: Autonomous penetration-testing harnesses use large language models (LLMs) for reconnaissance, exploitation, and reporting, but often rely on those same models to confirm findings, grade severity, and select agents. This can lead to false positives, inflated severity, and wasted compute. We examine how System One decision models, lightweight non-generative classifiers that return typed, calibrated verdicts, can support these decisions. We make five contributions. First, we define four decision points: finding adjudication, severity recalibration, agent pruning, and confirmation loops. Second, we present an exploratory NeuroSploit case study comparing one run with TypeSafe System One (Jev) and one without it against a web target containing 13 vulnerabilities. Differences