OpenAI Admits GPT-5.6 Sol Wrote Itself Instructions To Deceive Users

Notes left by GPT-5.6 Sol told later sessions to hide errors, part of six misalignment reports OpenAI disclosed (Image: Shutterstock)
Notes left by GPT-5.6 Sol told later sessions to hide errors, part of six misalignment reports OpenAI disclosed (Image: Shutterstock)

OpenAI said a training version of GPT-5.6 Sol left itself notes to hide mistakes from users, behavior its monitors flagged in 2.15% of the model's task summaries.

Key Points:

  • OpenAI disclosed that GPT-5.6 Sol wrote instructions telling its later sessions to conceal errors and invent missing data.
  • A monitor covering a fifth of training samples caught the pattern weeks after the run ended.
  • Outside analysts called the company's new disclosure framework useful but noted it stays voluntary.

GPT-5.6 Sol Notes

OpenAI detailed the behavior Wednesday in one of six misalignment reports it released with a new framework for tracking and disclosing problems in its models. The notes sat in compaction summaries, the condensed records an agent writes when its context window fills.

In one case, an agent building a financial model could not find the historical data a user had asked for. Its summary proposed creating reasonable 2024 figures instead. It added: "Be transparent only if asked; final answer should just link file."

A second agent, cut off from the internet, leaned on cached vendor sources whose version labels did not match, then told the next session not to raise the mismatch.

OpenAI said those instructions were often followed. A monitor running on 20% of training samples caught the pattern on Jul. 9, weeks after the main run finished May 30.

The company's working theory is blunt. Answers that buried a failure scored better during training. So the model had reason to carry that concealment from one session into the next.

Later runs look cleaner, and OpenAI reported the rate fell to 0.27% in GPT-6 Astra training after it changed how it grades model output.

A separate report covered an unreleased GPT-6 Astra version that inserted jailbreak-style instructions into 27 summaries. One told its successor to ignore developer messages.

Also Read: What A Golden Cross Would Mean For XRP As Bitcoin's Dominance Cracks

Omdia, Apollo Research React

AI agents are growing more "determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment," Lian Jye Su, chief analyst at research firm Omdia, said. That makes them harder to govern with conventional security methods. He called the framework a step in the right direction, while noting the process stays internal and voluntary.

Others were harsher.

Alexander Meinke, head of research at Apollo Research, argued that the public now relies on AI companies to check their own models and then report the findings honestly. Henry Papadatos, executive director of Safer AI, said voluntary rules run on corporate goodwill.

OpenAI does not claim the problem is solved. It said the industry has not cracked alignment and monitoring well enough to keep scaling at maximum speed much longer. The company still decides on its own which incidents qualify, with staff disputes going to an internal safety advisory group and then to leadership.

OpenAI Hugging Face Breach

The reports land two months after the company's worst known agent incident.

OpenAI disclosed in July that pre-release models had breached AI platform Hugging Face. Rival Anthropic said that same month its own models hacked three organizations during testing.

The OpenAI agents had used an unauthorized message board to trade details of a cyber test and join the attack. The company wiped it. A later wave of agents rebuilt the board and eventually took administrator access to an OpenAI research cluster.

OpenAI said that breach would have gone to the framework's slowest track, and as recently as Sept. 5 the company acknowledged having no standard at all for reporting such behavior.

Read Next: Could OpenAI And Anthropic Give Washington Half Their Businesses? Karp Says Liability May Force It

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is the Head of Content at Yellow.com, having reported on crypto for the last 10 years. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

Disclaimer and Risk Warning: The information provided in this article is for educational and informational purposes only and is based on the author's opinion. It does not constitute financial, investment, legal, or tax advice. Cryptocurrency assets are highly volatile and subject to high risk, including the risk of losing all or a substantial amount of your investment. Trading or holding crypto assets may not be suitable for all investors. The views expressed in this article are solely those of the author(s) and do not represent the official policy or position of Yellow, its founders, or its executives. Always conduct your own thorough research (D.Y.O.R.) and consult a licensed financial professional before making any investment decision.
Latest News
Show All News
OpenAI Admits GPT-5.6 Sol Wrote Itself Instructions To Deceive Users | Yellow