LATIDIA · Ciberseguridad
Comprender y mejorar la persistencia de la puerta trasera en la capacitación posterior del agente de LLM
arXiv:2610.07510v1 Announce Type: new Abstract: Los desarrolladores pueden crear agentes LLM adaptando modelos de terceros a través de un posentrenamiento benigno. Estudiamos una amenaza de cadena de suministro en la que un atacante suministra un modelo con un bac
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.07510v1 Announce Type: new Abstract: Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level reinforcement learning (RL). We observe that benign SFT substantially reduces attack success, but subsequent RL often preserves the residual behavior and sometimes even increases attack success. Our analysis of backdoor erosion during SFT identifies two factors that may favor survival: initial backdoor strength and gradient compatibility with benign