LATIDIA · Ciberseguridad
Huellas dactilares Modelos de lenguaje grandes multimodales
arXiv:2609.20457v1 Tipo de anuncio: nuevo Resumen: Si bien los modelos multimodales de lenguaje grande (MLLM) permiten una amplia gama de tareas de razonamiento imagen-texto, los incidentes recientes indican que son vulnerables a la implementación ilícita a
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.20457v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model provenance are typically confounded by shared language backbones in MLLMs and struggle to detect violations of distillation. To bridge this gap and safeguard model ownership, we present the first study on multimodal model fingerprinting. Inspired by recent findings that self-attention acts as a low-pass filter and that its low-frequency components are informative, we develop AttnPrint for white-box provenance. Specifically, we extract cross-modal attention distributions and isolate their low-frequency components to serve as model fingerprints.