LATIDIA · Ciberseguridad
Marcas de agua de comportamiento semántico: procedencia resistente a la falsificación y a la robustez de paráfrasis para agentes de LLM
arXiv: 2610.08668v1Announce Type: new Resumen: La marca de agua de comportamiento incrusta un identificador de propietario en las opciones de acción de alto nivel de un agente LLM, dando procedencia sin tocar los tokens de salida. Marcas de agua del agente anterior bre
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.08668v1 Announce Type: new Abstract: Behavioral watermarking embeds an owner identifier in an LLM agent's high-level action choices, giving provenance without touching output tokens. Prior agent watermarks break in two ways. First, all three prior schemes bind the watermark to the exact action symbol, so renaming a tool desynchronizes decoding even when the observation is untouched; in AgentMark's own robustness test, paraphrasing the observation alone drops bit-recovery to 16.8%. Second, every prior agent watermark studies only removal: none asks whether an adversary can forge a trajectory that verifies as someone else's, a question answered affirmatively for text watermarks (Jovanovi\'c et al., 2024). We present Semantic Behavioral Watermarking (SBW): watermarking over