LATIDIA · Ciberseguridad
ReproBench: Evaluación comparativa de agentes LLM sobre la reproducción de vulnerabilidades desde cero
arXiv: 2609.34450v1Announce Type: new Resumen: Los agentes del modelo de lenguaje grande (LLM) se evalúan cada vez más en tareas de ciberseguridad como la reproducción de vulnerabilidades, la explotación y los parches. Sin embargo, los ciberdelincuentes existentes
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.34450v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an executable binary. This setup bypasses the critical environment reconstruction step, leaving a fundamental question for real-world vulnerability analysis: can an agent autonomously reconstruct the required execution environment and reproduce a vulnerability entirely from scratch? To address this gap, we present ReproBench, an evidence-grounded benchmark designed to evaluate agent capabilities in end-to-end vulnerability reproduction starting from solely a CVE identifier. ReproBench decomposes the full reproduction workflow into