LATIDIA · Ciberseguridad
AuxMark: Defensa contra la destilación de agentes no autorizados a través de marcas de agua de comportamiento auxiliar
arXiv: 2609.34597v1Tipo de anuncio: nuevo Resumen: Los agentes de modelos de lenguaje grandes pueden adquirir capacidades complejas a través de la interacción de varios pasos y el uso de herramientas, pero sus trayectorias también se pueden recopilar ilegalmente para distinguir
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.34597v1 Announce Type: new Abstract: Large language model agents can acquire complex capabilities through multi-step interaction and tool use, but their trajectories can also be illegally collected to dis- till student agents. However, existing watermarking methods either do not fit the structured and interactive nature of agent environments or lack reliable effective- ness across tasks and model architectures. We introduce AuxMark, a behavioral watermarking framework for tracing unauthorized agent distillation. AuxMark dynamically inserts safe, non-essential auxiliary action into teacher trajectories, and stores the associated contexts as private evidence cards. To audit a suspicious student model, AuxMark constructs paired real and fake probes from these cards and applies a card-level sign