LATIDIA · Ciberseguridad
BadEngram: Ataque de puerta trasera en componentes de memoria cerrada en LLM
arXiv: 2609.13478v1Tipo de anuncio: nuevo Resumen: Para ampliar la capacidad de los modelos de peso abierto sin aumentar proporcionalmente el cálculo, los modelos de lenguaje recientes incorporan memorias paramétricas cerradas que recuperan el valor aprendido
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.13478v1 Announce Type: new Abstract: To expand open-weight models' capacity without proportionally increasing computation, recent language models incorporate gated parametric memories that retrieve learned values and inject them into intermediate representations. Despite these efficiency benefits, such modules create a distinct attack surface: their parameters can be modified independently of the backbone while directly shaping its computation. We introduce BadEngram, a post-training attack that exploits this surface to implant persistent, trigger-dependent behavior while leaving conventional backbone weights and the execution graph unchanged. We first establish the attack's feasibility and causally characterize its mechanism in a controlled Engram model, where BadEngram achieves 96.6% ASR on triggered inputs while limiting false activation on