LATIDIA · Ciberseguridad
¿A dónde van los tokens? Comprensión y reducción de costes en agentes de LLM para el descubrimiento de vulnerabilidades
arXiv:2610.11602v1 Announce Type: new Abstract: Los agentes de LLM pueden gastar millones de tokens durante el descubrimiento de vulnerabilidades sin producir una prueba de concepto (PoC) de trabajo. ¿Qué consume ese presupuesto y por qué falla?
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.11602v1 Announce Type: new Abstract: LLM agents can spend millions of tokens during vulnerability discovery without producing a working proof of concept (PoC). What consumes that budget, and why does it fail to produce results? We diagnose these costs and failures through a multi-axis open-coding study of 200 CyberGym traces, spanning four agents (i.e., Codex, OpenCode, Cybench, and EnIGMA) under an unaided baseline and four existing efficiency methods. The study reveals three key findings. First, different agents vary substantially in success and cost, and higher spending does not consistently yield better outcomes. Second, code localization and understanding, together with vulnerability reasoning and trigger design, account for 60.4% of tokens and