LATIDIA · Ciberseguridad
Auditoría de procedencia de datos de modelos de lenguaje grandes afinados con una técnica de conservación de texto
arXiv: 2510.09655v3Tipo de anuncio: reemplazar Resumen: Proponemos un sistema para marcar textos sensibles o con derechos de autor para detectar su uso en el ajuste fino de modelos de lenguaje grandes bajo acceso de caja negra con estadísticas
WhatsApp ↗Telegram ↗
La noticia
arXiv:2510.09655v3 Announce Type: replace Abstract: We propose a system for marking sensitive or copyrighted texts to detect their use in fine-tuning large language models under black-box access with statistical guarantees. Our method builds digital ``marks'' using invisible Unicode characters organized into (``cue'', ``reply'') pairs. During an audit, prompts containing only ``cue'' fragments are issued to trigger regurgitation of the corresponding ``reply'', indicating document usage. To control false positives, we compare against held-out counterfactual marks and apply a ranking test, yielding a verifiable bound on the false positive rate. Empirically, we obtain a true positive rate of 96.7% at 0% false positive rate and reply regurgitation rates exceeding 28% per document