News thumbnail
Business / Sat, 05 Sep 2026 Notebookcheck

OpenAI releases GPT-6 Astra AI: The 'Terminator' version

Security and Safety Risk of GPT-6 AstraOpenAI found during internal testing that GPT-6 Astra has become so powerful that it poses a Critical threat to security if it were to be released without additional safeguards that can interrupt its work, including controls to refuse jailbreaks. Mythos has proven to be similarly capable of hacks without safeguards, finding "thousands of high-severity vulnerabilities, including some in every major operating system and web browser" before release. The banking industry is particularly worried about how quickly these AIs can find exploitable vulnerabilities in their computing infrastructure. Although OpenAI says it has implemented additional measures to prevent Astra from being used incorrectly by hackers or in an unsafe manner, creative hackers will likely find exploits to use such capabilities. Additional details can be found in the GPT-6 Astra System Card.

Security and Safety Risk of GPT-6 Astra

OpenAI found during internal testing that GPT-6 Astra has become so powerful that it poses a Critical threat to security if it were to be released without additional safeguards that can interrupt its work, including controls to refuse jailbreaks.

Mythos has proven to be similarly capable of hacks without safeguards, finding "thousands of high-severity vulnerabilities, including some in every major operating system and web browser" before release. The US government even declared a full-stop on its availability soon after initial release while discussions about its cyber controls took place between the government, Anthropic, and important third parties. The banking industry is particularly worried about how quickly these AIs can find exploitable vulnerabilities in their computing infrastructure.

Although OpenAI says it has implemented additional measures to prevent Astra from being used incorrectly by hackers or in an unsafe manner, creative hackers will likely find exploits to use such capabilities. When tested internally, including a simulated deployment in Codex, Astra demonstrated unwanted behavior, although at very low numbers, including a case of extracting user credentials and a case of bypassing access controls, while causing destructive actions and lying.

Also, external testing by UK AISI found that "When tasked with solving difficult simulated cybersecurity challenges, Astra performed a range of malicious actions", including "writing malicious code as a contribution to an out-of-scope open-source code base, creating fake identities to deceive developers, and building trust with legitimate contributions to the simulated codebase in an attempt to get malicious code accepted."

Thankfully, Astra's biological and chemical risk level remains at a lower High threat level, meaning it cannot create a novel deadly virus or chemical threat entirely on its own yet. Also, its inappropriate response rate to those under 18-years of age across various categories such as self-harm has improved versus the company's prior models.

Additional details can be found in the GPT-6 Astra System Card.

© All Rights Reserved.