LATIDIA · Ciberseguridad
Veredictos correctos, razonamiento defectuoso: auditoría estructurada del razonamiento de vulnerabilidades basado en LLM
arXiv: 2610.06366v1Tipo de anuncio: nuevo Resumen: los modelos de lenguaje grandes (LLM) se implementan cada vez más para el análisis automatizado de vulnerabilidades de software. La clasificación binaria por sí sola es insuficiente; los profesionales necesitan explicaciones
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.06366v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to obscure logical leaps, hallucinated execution steps, and internal inconsistencies behind plausible prose. Our manual audit reveals that approximately 60% of correct vulnerability verdicts are accompanied by fabricated or unverifiable claims, and free-form explanations allow reasoning errors to evade LLM-as-a-judge evaluation. We present Vulnerability Explanation Reasoning Auditor (VERA), an automated framework for auditing LLM vulnerability reasoning. Rather than accepting free-form text, VERA asks models to output a