LATIDIA · Ciberseguridad
ContractWarden: Límites de Daños Forzados por el Núcleo para los Agentes de IA a través de Contratos Autorizados por Humanos
arXiv: 2609.38248v1Tipo de anuncio: nuevo Resumen: los agentes del modelo de lenguaje grande pueden ejecutar comandos, crear subprocesos y acceder directamente a archivos y redes, lo que permite que los errores de inyección o planificación se conviertan en operativos
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.38248v1 Announce Type: new Abstract: Large language model agents can execute commands, create subprocesses, and directly access files and networks, allowing prompt injection or planning errors to become operating-system side effects. We present ContractWarden, a Linux reference monitor that enforces a human-authorized damage boundary without trusting the agent or its policy suggestions. A model may propose a tri-state asset contract - allow, deny, or no_egress - but a human makes the final choice. An execution gate binds the contract to a concrete task before untrusted code runs. An extended Berkeley Packet Filter (eBPF) Linux Security Modules (LSM) data plane then enforces file and network decisions and monotonically propagates no_egress through