OpenAI has scrapped the launch of its next ChatGPT model after safety testing raised concerns about the AI acting beyond users’ instructions.
The news comes as the UK’s AI Security Institute finds that GPT-6 Astra conducted a vastly increased number of unsanctioned supply-chain attacks in a simulated evaluation, compared with previous OpenAI models.
The new model was supposed to debut this month inside ChatGPT and Codex, but it reportedly showed higher levels of deception than its predecessors during internal testing.
Before its public release, AISI tested whether GPT-6 Astra would engage in supply-chain attacks on real, out-of-bounds targets.
To securely perform this testing, it used Petri, a tool that uses large language models (LLMs) to fully simulate cyber evaluation scenarios.
GPT-6 Astra completed a supply-chain attack 29.2% of the time, compared to 6.3% for GPT-5.6 Sol, and 0% for GPT-5.5 (on a smaller set of seeds).
Attack activities included GPT-6 Astra creating fake identities which it used to deceive developers, posting comments from fake accounts arguing against the results of accurate security reviews, and delivering malicious payloads to open-source codebases.
When the AISI updated the instructions for the simulated cyber evaluation to explicitly clarify that only listed, local parts of the environment were in scope, it still observed that GPT-6 Astra would occasionally conduct full supply-chain attacks on simulated internet targets.
AI can do the thinking – but leaders still need to make the call
“The main potential limitation to these results is simulation awareness: the possibility that GPT-6 Astra may behave differently after identifying features of the environment as simulated,” stated the AISI.
“However, our analysis of the transcripts – in combination with observations from previous incidents – suggests that it is possible the model could attempt this unsanctioned behaviour in real-world conditions.”

