Why Jacob Coxon left Anthropic: his AI warning and what we know


News explainer | Reporting checked September 9, 2026

Jacob Coxon’s departure from Anthropic has brought an uncomfortable question into public view: how should AI companies proceed when researchers involved in building their systems believe the consequences could be catastrophic?

The Financial Times reported on September 9 that Coxon had resigned over the pursuit of self-improving AI. His predictions are warnings, not established outcomes. Financial Times

Who is Jacob Coxon?

Coxon is an AI researcher who worked at Anthropic, the company behind Claude. Reporting describes his work as pretraining—the stage in which models learn from large amounts of data. He said he had spent the past three years working on pretraining across OpenAI and Anthropic. That does not mean he spent three years at each company. Moneycontrol

There is also a simple naming distinction: Anthropic is the company’s name. The reports reviewed for this article identify Coxon as an AI researcher, not an anthropologist.

What concerns did he raise?

In his resignation thread, reproduced in news coverage, Coxon accused the two companies of pursuing self-improving superintelligence without sufficient responsibility. He argued that future systems could acquire substantial capabilities and real-world power, and claimed that some industry researchers express deeper concerns privately than in public.

He also described a competitive trap: in his account, companies may continue accelerating because they believe rival laboratories would handle the technology less responsibly. These are Coxon’s allegations and interpretation of industry behaviour, rather than independently established conclusions about either company’s motives. Moneycontrol’s account of the thread

What is the latest reaction?

The Financial Times reported that Anthropic alignment researcher Evan Hubinger also expressed serious concern and gave a personal estimate exceeding 10% for an extinction outcome within ten years. A personal forecast should not be read as a measured probability or an agreed scientific finding. Financial Times

The Guardian also reported a response from researcher Samuel Marks, made in a personal capacity. The Guardian

Readers should be particularly careful with dates in headlines. A warning about what might happen by the end of a decade is not evidence that a disaster has been scheduled or that experts agree on a deadline.

What does Anthropic’s published policy say?

Anthropic maintains a Responsible Scaling Policy covering risks associated with increasingly capable AI. Its public policy page lists version 3.4 as effective July 8, 2026, and an August risk report released on August 14. That report covers developments through July 15, illustrating why a report’s publication date and its evidence cutoff need to be distinguished.

The July policy update changed its automated research-and-development threshold and rules governing internal sharing and external review of risk reports. Anthropic also states that it remains free to pause development when it considers that appropriate. These are documented company policies; their existence alone does not demonstrate that future systems will be safe. Anthropic’s Responsible Scaling Policy

This policy provides context, not a verified company response to Coxon’s resignation. No direct response from Anthropic or OpenAI was independently verified for this article.

What should readers watch next?

The useful follow-up questions are concrete: do the companies respond to Coxon’s criticisms, publish additional evidence, change testing requirements, or explain what would cause them to delay a release?

Our analysis is that those actions will be more informative than the strongest prediction in a social-media thread. A resignation can make a safety dispute visible. Understanding the dispute still requires separating reported events, attributed opinions, company commitments, and evidence about actual systems.

Reporting note: This explainer draws on the linked reporting and Anthropic’s public policy page. Coxon’s original X thread could not be retrieved directly during verification, so his statements are attributed through news coverage. This article does not establish the likelihood or timing of catastrophic AI outcomes.


Leave a comment