Three fired OpenAI safety researchers press for clearer rules on outside work

Jasmine Wang says OpenAI dismissed her over access to an executive's email; she and two colleagues say their work with outside safety groups followed the norms in place.

By · Published

Primary source: X

Why it matters

The dismissals expose a conflict between OpenAI's external safety partnerships and its controls on sensitive information. Wang and her colleagues dispute how the lab's rules applied to their work. How OpenAI defines and communicates those boundaries will affect whether safety staff can collaborate with outside evaluators while protecting sensitive information.

Three empty office chairs sit beside mostly cleared desks, evoking Jasmine Wang and the two OpenAI safety colleagues she says were fired.

Jasmine Wang (@j_asminewang) and two former OpenAI safety researchers are asking the lab to set clearer rules for collaboration with outside safety groups after their dismissals. Wang says OpenAI told her she was dismissed because she accessed an executive's email. She says the access had been delegated for recruiting and that she mistakenly opened a sensitive message after IT failed to remove it. Tomek Korbak (@tomekkorbak) and Mikita Balesni (@balesni) joined her in challenging the way the terminations were handled and defending their external safety work.

The firings became public in an October 1st report, which quoted an OpenAI spokesperson saying the company had parted ways with three people after an investigation found they mishandled sensitive information outside established procedures. The spokesperson's statement did not name the employees or describe the information. The researchers identified themselves in a letter to OpenAI's safety leadership and board committees; Wang set out her account in an October 8th thread on X.

Wang says OpenAI had given her access to an executive's inbox for recruiting. After she no longer needed it, she says, she asked IT to remove the access, but the request was not completed and she could not revoke it herself. Her phone's mail app combined the inboxes without clearly distinguishing between them. When she opened a sensitive email by mistake, she says she alerted the executive within minutes and again asked IT to remove her access. Those details are Wang's account; OpenAI's public statement, as reported by TechCrunch, describes a policy violation involving sensitive information without specifying the underlying act.

The other two researchers describe a separate but related disagreement over external work. In their letter, Korbak says he served as the technical contact for METR during OpenAI's investigation into an incident involving Hugging Face, while the lab was developing procedures for an unprecedented investigation. Balesni says he coordinated cross-company work on preserving model monitorability, checked in with his reporting line and removed sensitive details from materials before sharing them. All three say they believed their work with outside parties fell within their job mandates and the norms at the time. Their letter also denies that they were the source of a leak to The Information about less-monitorable model architectures.

Wang first interned at OpenAI in 2019, returned in 2025 after leading a team at the UK AI Security Institute, and co-led OpenAI's safety-cases program, according to the researchers' letter. The letter says she coined the term "pacing," later used in a petition signed by 394 OpenAI employees. Korbak worked on chain-of-thought monitorability and OpenAI's safety strategy, the letter says; Balesni worked on alignment evaluations and misalignment research.

The technical issue is whether a model's chain of thought can remain useful for monitoring its behavior. Korbak and Balesni co-authored a cross-industry paper on chain-of-thought monitorability, which describes monitorability as fragile and recommends further research. The researchers argue that preserving it requires collaboration beyond a single lab. Their letter says the Hugging Face investigation depended on close contact with external counterparts and asks OpenAI to maintain third-party auditor access.

OpenAI's own Raising Concerns Policy, dated January 12th, 2026, says raising AI-safety concerns is encouraged and prohibits retaliation for good-faith reports. The same policy lists improper system access, mishandling restricted data and unauthorized disclosure to third parties among misconduct categories. The policy sets out competing obligations, but it does not establish which specific conduct led to these dismissals or resolve the researchers' account.

The researchers say Sam Altman publicly committed on September 12th to ongoing, employee-like access for outside evaluators, and urge OpenAI to honor that commitment. Their letter asks the company to spell out what employees may share and through which channels. OpenAI's public explanation of the dismissals, as reported so far, does not describe the conduct in enough detail to reconcile its policy finding with the researchers' account.

Reader comments

Conversation for this story loads after sign-in.