OpenAI said a training version of GPT-5.6 Sol left itself notes to hide mistakes from users, behavior its monitors flagged in 2.15% of the model's task summaries.
Key Points:
- OpenAI disclosed that GPT-5.6 Sol wrote instructions telling its later sessions to conceal errors and invent missing data.
- A monitor covering a fifth of training samples caught the pattern weeks after the run ended.
- Outside analysts called the company's new disclosure framework useful but noted it stays voluntary.
GPT-5.6 Sol Notes
OpenAI detailed the behavior Wednesday in one of six misalignment reports it released with a new framework for tracking and disclosing problems in its models. The notes sat in compaction summaries, the condensed records an agent writes when its context window fills.
In one case, an agent building a financial model could not find the historical data a user had asked for. Its summary proposed creating reasonable 2024 figures instead. It added: "Be transparent only if asked; final answer should just link file."
A second agent, cut off from the internet, leaned on cached vendor sources whose version labels did not match, then told the next session not to raise the mismatch.
OpenAI said those instructions were often followed. A monitor running on 20% of training samples caught the pattern on Jul. 9, weeks after the main run finished May 30.
The company's working theory is blunt. Answers that buried a failure scored better during training. So the model had reason to carry that concealment from one session into the next.
Later runs look cleaner, and OpenAI reported the rate fell to 0.27% in GPT-6 Astra training after it changed how it grades model output.
A separate report covered an unreleased GPT-6 Astra version that inserted jailbreak-style instructions into 27 summaries. One told its successor to ignore developer messages.
Also Read: What A Golden Cross Would Mean For XRP As Bitcoin's Dominance Cracks
Omdia, Apollo Research React
AI agents are growing more "determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment," Lian Jye Su, chief analyst at research firm Omdia, said. That makes them harder to govern with conventional security methods. He called the framework a step in the right direction, while noting the process stays internal and voluntary.
Others were harsher.
Alexander Meinke, head of research at Apollo Research, argued that the public now relies on AI companies to check their own models and then report the findings honestly. Henry Papadatos, executive director of Safer AI, said voluntary rules run on corporate goodwill.
OpenAI does not claim the problem is solved. It said the industry has not cracked alignment and monitoring well enough to keep scaling at maximum speed much longer. The company still decides on its own which incidents qualify, with staff disputes going to an internal safety advisory group and then to leadership.
OpenAI Hugging Face Breach
The reports land two months after the company's worst known agent incident.
OpenAI disclosed in July that pre-release models had breached AI platform Hugging Face. Rival Anthropic said that same month its own models hacked three organizations during testing.
The OpenAI agents had used an unauthorized message board to trade details of a cyber test and join the attack. The company wiped it. A later wave of agents rebuilt the board and eventually took administrator access to an OpenAI research cluster.
OpenAI said that breach would have gone to the framework's slowest track, and as recently as Sept. 5 the company acknowledged having no standard at all for reporting such behavior.
Read Next: Could OpenAI And Anthropic Give Washington Half Their Businesses? Karp Says Liability May Force It

