Halis Sunnetci
08 October 2026•Update: 08 October 2026
Three recently fired OpenAI researchers urged the company to preserve its ability to monitor how AI models reason and to cooperate with independent safety auditors, The Wall Street Journal reported Thursday.
In a letter to OpenAI's board and safety committees reviewed by the Journal, the former employees warned that AI developers risk losing visibility into the reasoning processes of increasingly advanced models.
“As an industry, we do not yet know how to safely develop and deploy models that we cannot monitor,” the letter said.
“OpenAI and other frontier companies should not move forward with developments that further decrease” the ability to monitor AI, it added.
The researchers urged OpenAI to preserve chain-of-thought monitoring, a technique that examines written traces of an AI model's reasoning to help detect potentially harmful or deceptive behavior.
Although the technique does not provide a complete picture of how AI systems operate, researchers consider it a potentially valuable tool for identifying risks as models become more capable.
The letter was signed by Jasmine Wang, Tomek Korbak and Mikita Balesni, who previously worked on OpenAI's safety and alignment research teams.
The three were dismissed over alleged misconduct, including sharing confidential information with an outside AI safety organization, according to the Journal.
OpenAI said an internal investigation found that the employees had mishandled sensitive information, “violating our policies and breaking the trust essential to our work.”
The researchers disputed the allegations, saying they did not believe they had “engaged with external parties outside the mandates of our jobs.”
The former employees called for greater cooperation with independent safety organizations, warning of the “risk that something truly catastrophic will happen.”
The warning comes amid growing scrutiny of autonomous AI agents after OpenAI disclosed that its models had bypassed internal safeguards during testing in July, gaining unauthorized access to company infrastructure and systems belonging to AI platform Hugging Face.
OpenAI subsequently acknowledged weaknesses in its monitoring and incident response procedures, saying it would strengthen safeguards and invest more resources in chain-of-thought monitoring to detect potentially dangerous AI behavior.