LATIDIA · Ciberseguridad
SkillPoison: envenenamiento progresivo de habilidades a través de experiencias exitosas
arXiv: 2610.07645v1Tipo de anuncio: nuevo Resumen: Los agentes de LLM que se mejoran a sí mismos destilan cada vez más experiencias exitosas en habilidades persistentes y reutilizables. Los métodos de ataque de habilidades existentes corrompen este proceso de aprendizaje por inje
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.07645v1 Announce Type: new Abstract: Self-improving LLM agents increasingly distill successful experiences into persistent, reusable skills. Existing skill attack methods corrupt this learning pipeline by injecting malicious triggers, behaviors, or false facts into individual experiences or extracted skills. However, such attacks are easily detected, and the injected malicious behaviors often fail to accumulate as persistent skills. In this paper, we show that skill poisoning can arise even from verified successful experiences, without making any individual trajectory malicious. Based on this insight, we propose SkillPoison, a novel framework that progressively poisons skill via successful experiences. SkillPoison first constructs a set of successful experiences that reinforce a target behavior, and then removes