LATIDIA · Ciberseguridad
MarkSec: Evaluación consciente de la capacidad de los ataques adversarios contra las marcas de agua de LLM
arXiv: 2609.16681v1Tipo de anuncio: nuevo Resumen: la marca de agua LLM ayuda a rastrear el origen del texto generado, pero se enfrenta a ataques de robo que recuperan información de marca de agua, ataques de depuración que eliminan señales de marca de agua, un
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.16681v1 Announce Type: new Abstract: LLM watermarking helps trace the origin of generated text, but faces stealing attacks that recover watermark information, scrubbing attacks that remove watermark signals, and spoofing attacks that forge text accepted as watermarked. These attacks are often studied in isolation, leaving their connections unclear. Evaluations also often lack shared detector calibration, metric definitions, and reporting protocols. Moreover, measuring attack success and text quality separately makes it difficult to identify attacks that are both effective and quality-preserving. We propose MarkSec, a general framework that unifies analyses of stealing, scrubbing, and spoofing. We evaluate attacks under a common reporting protocol and introduce a quality-constrained attack success metric to