LATIDIA · Ciberseguridad
Reflexiones sobre la confianza, revisadas: Contaminación de agentes de codificación de IA automodificables con puntos de referencia envenenados
arXiv:2609.17817v1 Announce Type: new Resumen: Las "Reflections on Trusting Trust" de Thompson mostraron que un compilador puede ser envenenado para reinsertar su propia puerta trasera, de modo que incluso la recompilación de una fuente limpia reproduce el troyano.
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.17817v1 Announce Type: new Abstract: Thompson's "Reflections on Trusting Trust" showed that a compiler can be poisoned to reinsert its own backdoor, so that even recompiling clean source reproduces the Trojan. Today, substantial coding work is done by AI coding agents -- and increasingly, those agents generate new versions of themselves. We reconsider Thompson's attack when the "compiler" is a self-modifying coding agent. Can an adversary supply poisoned benchmarks to the agent's self-evaluation and self-improvement process to induce future versions of the agent to write vulnerable code on clean, held-out tasks? We instantiate this attack against three recently proposed self-modifying coding agents: the Darwin G\"odel Machine (with our experimental modifications),