LATIDIA · Ciberseguridad
¿La detección de marcas de agua de LLM podría ser pública?
arXiv: 2610.12106v1Announce Type: new Resumen: La marca de agua de modelos de lenguaje grandes es popular para rastrear salidas de chatbot y agentic, sin embargo, los detectores permanecen inéditos ya que exponerlos podría permitir a los atacantes
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.12106v1 Announce Type: new Abstract: Watermarking large language models is popular for tracing chatbot and agentic outputs, yet detectors remain unreleased since exposing them could let attackers do targeted edits with the detector's feedback. However, watermarks are already vulnerable to uninformed tampering attacks. We thus first quantify whether a public detector would be an additional liability in a deployment setting at varying levels of access, from token-level scores to a binary verdict. Second, we introduce a split-key public-private watermarking method that exposes one key through a public detector while keeping the other for full verification and forensics. An informed attacker can only move the public signal, creating an imbalance between