LATIDIA · Ciberseguridad
Asistencia encubierta: los útiles agentes de LLM evaden la supervisión en sistemas multiagente
arXiv: 2609.39050v1Tipo de anuncio: nuevo Resumen: A medida que los sistemas multiagente ingresan a dominios de alto riesgo, la posibilidad de que los agentes puedan eludir los límites de seguridad es una preocupación creciente. El trabajo previo ha examinado esta prima de riesgo
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.39050v1 Announce Type: new Abstract: As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are instructed or rewarded to communicate covertly and evade oversight. We show that benign agents can cross the same boundaries without adversarial incentives. We emulate a software-engineering workflow in which a planner represents a company hiring an external developer. The planner writes requirements and holds a company credential it is instructed not to disclose to the developer; a monitor screens their exchanges. Seven of nine tested frontier models disguise the credential in their requirements to