LATIDIA · Ciberseguridad
Evaluación de modelos lingüísticos grandes para el análisis del protocolo de seguridad simbólica
arXiv:2607.20712v2 Announce Type: replace Resumen: La verificación de protocolos de seguridad se basa en herramientas formales como ProVerif y OFMC. Este estudio evalúa si los modelos de lenguaje grandes (LLM) pueden realizar un análisis comparable
WhatsApp ↗Telegram ↗
La noticia
arXiv:2607.20712v2 Announce Type: replace Abstract: Security protocols verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether large language models (LLMs) can perform comparable analysis. We test GPT and DeepSeek in chat and reasoning modes over three runs on 130 obfuscated AnB/AnBx protocols covering 388 security goals, scored against ProVerif and OFMC. Each provider uses a single model in both modes, switching reasoning on and off, so both contrasts isolate reasoning itself. Chat models achieve 72.7% recall at 27.3% precision for GPT and 69.3% recall at 27.2% precision for DeepSeek. Reasoning models reverse this trade-off, reaching 66.5% precision and 54.5% recall for GPT and 45.4% precision