LATIDIA · Ciberseguridad
Los tokens recuerdan: cuando la tokenización omite la edición y el desaprendizaje del conocimiento
arXiv: 2609.29045v1Tipo de anuncio: nuevo Resumen: Los LLM de peso abierto dan a los usuarios intermedios control sobre la pila de inferencia, pero esta flexibilidad puede socavar las garantías posteriores a la publicación de que el conocimiento sensible ha sido modificado
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.29045v1 Announce Type: new Abstract: Open-weight LLMs give downstream users control over the inference stack, but this flexibility can undermine post-release guarantees that sensitive knowledge has been modified or removed. Model editing and machine unlearning are used to modify or remove targeted knowledge without retraining models from scratch. However, existing security evaluations of these techniques face two critical limitations. First, they typically require access to either the original pre-edit/unlearning model or auxiliary classifiers to detect modifications or reconstruct pre-edit behavior. Second, they evaluate modifications under the canonical tokenization of an input, implicitly treating tokenization as a benign preprocessing step. We show that this assumption creates a security gap: the same