LATIDIA · Ciberseguridad
Sobre la fiabilidad de los puntos de referencia de parches de vulnerabilidad basados en LLM
arXiv: 2610.10150v1Tipo de anuncio: nuevo Resumen: Los modelos de lenguaje grandes (LLM) han mostrado un gran potencial para el parcheo automatizado de vulnerabilidades, pero los puntos de referencia actuales pueden distorsionar sustancialmente el rendimiento informado. Drawin
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.10150v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential for automated vulnerability patching, but current benchmarks can substantially distort reported performance. Drawing on extensive experience developing, running, and stress-testing such frameworks, we identify under-examined pitfalls across three dimensions: (1) agent-level factors, where prompting, tool availability, and detailed instructions can raise success rates without improving developer-aligned patch quality; (2) framework-level factors, where permission errors, infrastructure bugs, and timeout handling can silently suppress or inflate performance; and (3) dataset-level factors, where bug reports and single proof-of-concept (PoC) tests fail to capture whether patches address root causes or follow developer intent. We curate 112 historical bugs from 84 open-source