LATIDIA · Ciberseguridad
Extracción de secretos prácticos contra LLM de caja negra
arXiv: 2609.36941v1Tipo de anuncio: nuevo Resumen: Los modelos de lenguaje grandes (LLM) potencian cada vez más a los agentes de codificación autónomos como Codex y Claude Code, sin embargo, sus corpus de capacitación pueden contener credenciales confidenciales expo
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.36941v1 Announce Type: new Abstract: Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extraction audits, however, largely assume access to model weights or token probabilities. In this work, we present a black-box secret extraction framework for commercial, API-based LLMs under output-only access. It comprises (i) \emph{Cross-Validated Secret Knowledge Distillation}, which uses semantics-preserving prompt variants, response cross-validation, and provider-specific format filtering to distill secret-relevant behavior into a local white-box proxy; and (ii) \emph{Proxy-Guided Secret Extraction and Candidate Filtering},