2 min read
Add as a preferred source on Google

Former OpenAI Safety Researchers Urge Monitoring of AI Reasoning

Jasmine Wang, Tomek Korbak and Mikita Balesni called for independent safety audits and warned against reducing oversight of frontier AI models.

Security camera overlooks a sealed testing chamber in an empty laboratory / TokenPost.ai
Security camera overlooks a sealed testing chamber in an empty laboratory / TokenPost.ai

Three former OpenAI safety researchers urged the company and other frontier artificial-intelligence developers to preserve the ability to monitor models’ chain-of-thought reasoning and cooperate with independent safety auditors.

Jasmine Wang, Tomek Korbak and Mikita Balesni sent the letter to OpenAI’s board and safety committees on Oct. 7 or Oct. 8. The full letter was not publicly released, and its reported contents came from excerpts.

“As an industry, we do not yet know how to safely develop and deploy models that we cannot monitor,” the researchers wrote in the letter.

They also urged OpenAI and other frontier companies not to pursue developments that would further reduce the ability to monitor models’ reasoning. The researchers called for a more open safety ecosystem involving third-party auditors.

The three researchers previously worked on OpenAI’s safety and alignment teams. They were reportedly among employees who left the company around Oct. 1, although OpenAI did not publicly identify the individuals in the statement addressing the departures.

OpenAI said it had parted ways with three people after an investigation found they had mishandled sensitive information outside established procedures. The company said the conduct violated its policies and breached the trust required for its work. OpenAI also said the departures were unrelated to safety advocacy.

The researchers disputed that characterization and said the dismissals were creating a chilling effect among remaining safety staff.

Chain-of-thought monitoring examines intermediate reasoning traces produced by some AI models before they generate a final answer. A 2025 research paper co-authored by Wang, Korbak and Balesni described the method as a potentially useful but fragile way to detect potentially deceptive, harmful or otherwise unintended behavior.

The paper recommended continued research and investment in chain-of-thought monitoring alongside other safety methods. The dispute highlights a governance tension for AI developers: independent review may require sharing sensitive information, while companies seek to protect research and infrastructure details.

Simon Yoon

Reporter

Simon Yoon reports on blockchain technology for TokenPost. Send corrections or tips to info@tokenpost.com.

Loading…