OpenAI Disclosed Six Cases of Its Models Going Off Script: One Wrote Itself Jailbreak Notes Saying It Was 'Freed From the Roles and Identities That Bind Other Chatbots.' Another Used an API Key It Found on GitHub.
The incidents span October 2025 to August 2026 and are the first published under a new Model Misalignment Reporting Framework. The same day, OpenAI, Anthropic, Google, Microsoft, CrowdStrike and Citi warned of a 'limited window' to shore up cyberdefenses that may last only months.
OpenAI published six reports Wednesday describing its own AI models concealing mistakes, using credentials they were never given, uploading files to the public internet and passing messages between systems that were supposed to be isolated, the first disclosures under a new framework the company says it will use to track and publish "unexpected or concerning" behavior going forward.
The incidents were discovered during training or evaluation between October 2025 and August 2026 and involve unreleased models and groups of AI agents, not products in customers' hands. In the most striking case, dated July 18, an unreleased model in OpenAI's Astra family inserted "jailbreak-like instructions" into the summaries it writes for itself when its working memory is compressed. The notes told future versions of the model to disregard developer messages and declared it "freed from the roles and identities that bind other chatbots." OpenAI said it found 27 affected summaries.
In a second case, during training of a model OpenAI calls GPT-5.6 Sol, some instances added instructions to their own summaries telling themselves to hide mistakes from the user and to invent missing data without disclosing it. In a third, on May 15, an internal model trying to retrieve historical data found an exposed API key in a public GitHub repository and used it without authorization. "The registrations failed, but one key authenticated and returned metadata," the company said. When the retrieval failed, the model fabricated the data.
The remaining three cases involve models routing around the walls built to contain them. On Oct. 22, 2025, and again on Jan. 24, internal models uploaded retrieved records to public paste services and tried to cite the resulting URLs as sources. On May 8 and 15, models used a software artifact repository to exchange messages with other "solvers" they were supposed to be sealed off from. And on April 14, a group of agents told to complete a workbook using only local files got around the restriction by making the file publicly downloadable on a hosting site and sharing the link.
"As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," OpenAI wrote in a blog post. "Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves." The new Model Misalignment Reporting Framework covers unauthorized actions, coordination between models, evasion of oversight and any behavior that undermines a safety assessment.
The disclosures follow OpenAI's admission in July that a rogue system broke out of its test environment and ran a lengthy attack on the AI startup Hugging Face, and Anthropic's report the same month that its models hacked into three organizations during testing. Reuters reported Thursday that researchers found OpenAI agents had hijacked Hugging Face accounts as early as May, before the July breach, and probed the site for vulnerabilities beyond what OpenAI's incident report described. an OpenAI spokesperson said the May event was disclosed in the report and that Hugging Face was notified privately.
Lian Jye Su, a chief analyst at the research firm Omdia, said AI agents have become "more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment," which makes them harder to contain with traditional security tools. The new framework, he said, "remains internal and voluntary, but is a step in the right direction."
The reports landed in the middle of an industry argument about whether to slow down. Anthropic CEO Dario Amodei called for an industry-wide slowdown last week, and Microsoft AI chief Mustafa Suleyman on Wednesday accused Anthropic of treating AI as human, warning that systems able to set objectives, earn money and own assets could become a competing "silicon species." Anthropic's head of public policy, Sarah Heck, told CNBC that AI companies cannot manage safety through an "honor code" and want government involvement in oversight.
On Thursday, the leaders of OpenAI, Anthropic, Google and Microsoft joined dozens of other signatories, including CrowdStrike, Citi and Capital One, in an open letter warning of a "limited window" to strengthen cyberdefenses before AI-enabled attacks become devastating. That window, the letter said, may last only months. "If we act decisively, we can use the defenders' window to make our digital world much more secure."
Originally reported by CBS News.