Three researchers fired by OpenAI urged its board in an Oct. 7 letter to halt work that makes AI harder to monitor, and a company memo backed their advice.
Key Points:
- Three researchers dismissed by OpenAI sent its board a letter on Oct. 7 urging a stop to work that reduces AI monitoring.
- OpenAI said the firings were unrelated to safety concerns, and an internal memo strongly agreed with the letter's recommendations.
- The system card for GPT-6 Astra says the OpenAI model is harder to monitor than GPT-5.6 Sol.
OpenAI Board Letter Details
Tomek Korbak, Mikita Balesni and Jasmine Wang, who worked on safety and alignment at the company, addressed the letter to OpenAI's board and its safety committees, according to excerpts published Wednesday.
The full text has not been released. The researchers wrote that "as an industry, we do not yet know how to safely develop and deploy models that we cannot monitor."
They said OpenAI and its rivals should not pursue work that reduces that ability further. The letter reportedly also asks OpenAI to work with third-party auditors and support an open and transparent safety ecosystem, and it says the firings are "chilling those who remain at OpenAI."
Also Read: CrowdStrike Suspects A 26-Year-Old In China Hacked Korean Banks With Claude Code
OpenAI Memo And Firings
An OpenAI spokesperson said the dismissals "were not about raising safety concerns or speaking out." The spokesperson shared part of an internal memo, sent Wednesday by a research leader, that "strongly agreed" with the letter's recommendations. OpenAI said Oct. 1 that an investigation found three employees had mishandled sensitive information and broken company policy, in a case involving an outside group that evaluates AI models.
OpenAI has not publicly named the researchers or that organization.
GPT-6 Astra Monitoring Concerns
Chain of thought is the step-by-step reasoning a model writes out before it answers, and safety teams read it for signs that a system intends to misbehave.
OpenAI's system card for GPT-6 Astra, published Sept. 3, states that the model is harder to monitor this way than GPT-5.6 Sol and more capable of controlling its own chain of thought. Astra-class models "could evade our CoT monitors under adversarial conditions," the card says.
Those findings come largely from tests that instructed the model to evade monitoring, and OpenAI reported no evidence of Astra hiding its reasoning inside unrelated text.
The company called preserving that monitoring a core research goal and wrote that it will not accept further loss beyond a limit without new ways to show its models are aligned.
The three researchers have a record on the subject. Korbak and Balesni led a July 2025 paper, co-authored by Wang and researchers from OpenAI, Google DeepMind and Anthropic, that called chain-of-thought monitoring a new and fragile opportunity for AI safety. OpenAI's Astra system card cites that paper, which urged developers to weigh how their design decisions affect the ability to monitor a model's reasoning.
Read Next: AI Agents Without The Cloud: Microsoft, Nvidia Ready A 128GB Surface Laptop

