LATIDIA · Ciberseguridad
El desafío de identificar el origen de los modelos de lenguaje grandes de caja negra
arXiv:2503.04332v2 Tipo de anuncio: reemplazar Resumen: El tremendo potencial comercial de los modelos de lenguaje grandes (LLM) ha aumentado las preocupaciones sobre su uso no autorizado. Para abordar esto, nos centramos en la tarea de identificar
WhatsApp ↗Telegram ↗
La noticia
arXiv:2503.04332v2 Announce Type: replace Abstract: The tremendous commercial potential of large language models (LLMs) has heightened concerns over their unauthorized use. To address this, we focus on the task of identifying the origin of black-box LLMs. We further propose PlugAE, an effective and efficient identification method that proactively leverages LLM-specific adversarial embeddings and allows users to customize copyright tokens on a targeted query set. Extensive experiments demonstrate that PlugAE outperforms both state-of-the-art model watermarking and fingerprinting methods in accuracy and robustness. We further analyze its stealthiness and reliability from three complementary perspectives and conduct ablation studies under various configurations, confirming its practicality for real-world misuse detection.