LATIDIA · Ciberseguridad
Detección semiautomática de brechas en el conocimiento de seguridad de LLM
arXiv: 2607.18496v4Tipo de anuncio: reemplazar Resumen: Los modelos de lenguaje grandes (LLM) se utilizan cada vez más para una variedad de tareas de seguridad centradas en el software, el hardware y las personas. En consecuencia, el rendimiento de LLM en tareas de seguridad
WhatsApp ↗Telegram ↗
La noticia
arXiv:2607.18496v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM performance on security tasks is an active area of measurement and research, often with a focus on identifying areas in which LLM security "knowledge" may be insufficient. Popular strategies for identifying LLM security knowledge gaps include building corpora of challenge questions or task benchmarks, strategies that require substantial manual work and security expertise to design and execute. We introduce a partially-automated method for assessing LLM knowledge of a security area. The method uses authoritative information from Consumer Protection Agencies (CPAs) to identify instability in LLM responses