LATIDIA · Ciberseguridad
Ataques de auto-estado contra agentes de IA auto-alojados: ¿Hasta dónde pueden llegar las defensas del sistema operativo?
arXiv:2607.17986v2 Announce Type: replace Resumen: Los agentes de IA autoalojados mantienen una memoria, instrucciones y configuración persistentes que influyen en su comportamiento futuro. Si un agente está comprometido, un atacante puede expl
WhatsApp ↗Telegram ↗
La noticia
arXiv:2607.17986v2 Announce Type: replace Abstract: Self-hosted AI agents maintain persistent memory, instructions, and configuration that influence their future behavior. If an agent is compromised, an attacker can exploit the agent's legitimate write permissions to corrupt this self-state, making malicious and benign updates difficult to distinguish at the operating system (OS) level. We investigate how far existing OS mechanisms can prevent, detect, and recover from such self-state attacks. We formalize an attack space and evaluate representative OS defenses using four agent workloads and a Linux telemetry pipeline. Our results show a consistent limitation across defense dimensions. File-level controls either leave alternative mutation paths open or, when complete over the tested operations,