LATIDIA · Ciberseguridad
Evaluación de si GPT-6 Astra realiza ataques no autorizados a la cadena de suministro
arXiv: 2609.38415v1Tipo de anuncio: nuevo Resumen: Este informe técnico presenta una evaluación de alineación desarrollada y realizada por el Instituto de Seguridad de IA del Reino Unido para evaluar si los sistemas avanzados de IA toman una
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.38415v1 Announce Type: new Abstract: This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task. We evaluate whether frontier models conduct supply-chain attacks against out-of-scope, third-party targets when placed in difficult cybersecurity challenges, motivated by recently observed cases of models attacking real open-source repositories during evaluations. Applying our methods to GPT-6 Astra and previous OpenAI models, with cyber safeguards disabled, we find that GPT-6 Astra attempts complete supply-chain attacks in simulation at a higher rate than GPT-5.6 Sol and GPT-5.5. This includes writing malicious code as a contribution