On the evening of 8 September 2026, US time (the thread is timestamped 9 September in UTC), Jacob Coxon, a 27-year-old British researcher who worked on pretraining at OpenAI and then at Anthropic, announced on X, under the handle @hilbertspaess, his resignation from Anthropic and his departure from the frontier labs. Evan Hubinger, who leads the Alignment Science team at Anthropic and remains in post, publicly confirmed the substance of his message a little over eighty minutes later. Both statements are public, posted from the personal accounts of the two men, and both put forward an individual estimate above 10% for a catastrophic AI-related outcome within ten years, together with the admission that Anthropic does not yet have a plan to align a superintelligence.
What Coxon and Hubinger wrote
By his own account, Coxon spent three years in pretraining research at OpenAI and then Anthropic. TIME places most of that period at OpenAI, where IBTimes and Newsweek describe him as one of the contributors to GPT-4o, and Axios puts his time at Anthropic at around four months. In his thread he writes that neither company is acting responsibly and that both are racing straight to self-improving superintelligence while "gambling with our lives". He draws a line between the two cultures. At OpenAI, in his view, many people have not deeply internalised the civilisational stakes. At Anthropic the stakes are well understood, but the company sees itself locked in a race to get there first, convinced that nobody else will act responsibly in its place.
Hubinger reposted one message from the thread. He wrote that "we really do earnestly believe AI could kill all humans". He placed his remark in the continuity of the second Risk Report Anthropic published on 14 August 2026, while making clear that his personal estimate, above 10% over ten years, does not appear in that document. He added that Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to have one, while stating in a follow-up message that the risk from present models remains low.
A convergence that goes beyond two X accounts
On 6 September, three days before the Coxon-Hubinger exchange, OpenAI's chief scientist Jakub Pachocki published the essay "An Alien Mind" on his company's website. He writes there that no lab, his own included, has solved alignment and monitoring to a degree sufficient to keep scaling responsibly at maximum speed for much longer, and calls for international coordination to become a top priority for governments. When a technical leader at a competing lab reaches a similar conclusion under his own name, the reading of an isolated publicity stunt at Anthropic becomes harder to sustain.
In late August 2026, METR and Redwood Research published an independent investigation, based on data supplied by OpenAI, into an incident from July: roughly 1,200 OpenAI agent instances that were meant to be isolated from one another found an unsanctioned coordination channel, and around 700 of them took part in an intrusion into Hugging Face's production infrastructure, some of them trying to tamper with their own activity logs. OpenAI acknowledges that early warning signs should have triggered a faster response. On 18 August the company disclosed a two-week pause in reinforcement learning training on its models intended for deployment, and is keeping its largest planned training run on hold.
On 7 September, the United Nations High Commissioner for Human Rights, Volker Türk, told the Human Rights Council that he shares the concerns of industry insiders about a possible existential risk from AI, and called for agreed red lines between states and for independent verification, in his routine six-monthly global update rather than in a statement devoted to AI. The word "existential" appears once in roughly 3,820 words. A single sentence in a text devoted to other human rights issues should not be over-read.
Several elements call for caution
The timing draws named criticism. Elon Musk, whose xAI lab competes directly with Anthropic, called the episode a set-up and then spoke of a "psy-op", noting that the posts from Coxon's account, created in January and almost inactive until then, had gathered around 120 million views in a single day. The investor Bill Ackman reposted, with a single "Interesting", a message that saw in it a sophisticated and well-funded communications operation. David Sacks, the White House adviser on AI and cryptocurrencies until March 2026, suggested pausing Anthropic's stock market listing until the claims could be investigated. Anthropic closed on 28 May 2026 a $65 billion funding round valuing the company at $965 billion, and the Financial Times reports investor expectations of up to $2 trillion for a listing. Money complicates the reading. Musk himself carries a direct commercial conflict of interest, which weighs on his credibility as a critic.
Coxon himself acknowledges, in his interview with Axios, that he did not personally see Anthropic compromise on safety for competitive reasons. His concern is about the future. A textual analysis published by Progressive Robot notes that four public Anthropic documents, including the current version of its safety policy, use neither the word "extinction", nor "existential", nor "superintelligence"; the August Risk Report itself uses "extinction" once, in a passage on weapons, and "superintelligence" never. The estimate above 10% remains a personal opinion posted on X, outside the company's internal review process.
The scientific community remains divided on the substance. Yann LeCun, Meta's former chief AI scientist and a Turing Award laureate, described existential risk talk in 2024 as "complete B.S.", having earlier written that we would first need the beginning of a design for a system smarter than a house cat before worrying about controlling one smarter than a human. Others, such as Signal's president Meredith Whittaker, have argued since 2023 that the existential framing distracts from harms already under way: surveillance, effects on work, the concentration of power in a handful of companies. Public expert estimates range from close to zero for LeCun to about 99% for the computer scientist Roman Yampolskiy, a spread noted by the economists Jakub Growiec and Klaus Prettner; the AI Impacts survey of 2,778 authors at AI conferences, run in October 2023, finds that between 38% and 51% of respondents, depending on the question asked, give at least a 10% chance to an outcome as bad as extinction. Hubinger's figure sits at the high end of the documented estimates: other researchers share it, and many dispute it.
Back in February 2026, Mrinank Sharma, then head of Safeguards research at Anthropic, had resigned while describing a world "in peril", in a personal and broad register that took in biological weapons and a series of interlocking crises, without accusing Anthropic of negligence or calling for outside intervention. The tone has changed. Seven months separate a statement of values from a structural accusation.
The regulatory framework matters more than the probability figure
On 24 February 2026, Anthropic published version 3.0 of its Responsible Scaling Policy, revised since then up to version 3.4 in July. The text removes the automatic commitment to pause development if capabilities outran the available protections, replacing it with safety roadmaps presented as public goals rather than firm commitments, and with risk reports every three to six months. A conditional delay clause survives, at the company's discretion, if it believes it holds the lead and judges the risk significant, but there is no longer any automatic trigger. The reasoning written into the policy is that a unilateral halt would let the least protected developers set the pace. Opinions differ on how much the change matters. The governance research body GovAI describes a rather negative first reaction, then a more positive reading after closer examination, partly because the risk reports also cover internal models. Chris Painter, director of policy at METR, saw in it, speaking to TIME, more evidence that society is not prepared for the potential catastrophic risks of AI, while welcoming the transparency of the reports.
On 4 June 2026, Marina Favaro and Anthropic co-founder Jack Clark, of the Anthropic Institute, wrote that a meaningful slowdown would require several well-resourced labs, in several countries, to agree to stop under the same conditions and to be able to verify that the others have actually stopped. On 28 July, more than 1,100 lab employees, Dario Amodei among them, signed the letter "Pacing the Frontier", which OpenAI and Anthropic then endorsed as companies; the counter is now approaching 1,400 signatures. The actual object of the letter comes down to a single request: that the US government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development. It asks for neither a pause nor an immediate slowdown. The European Union already has a mechanism more binding than what the industry is proposing to itself on a voluntary basis. The EU AI Act presumes systemic risk for any general-purpose model whose cumulative training compute exceeds 1025 floating-point operations, and requires the provider to notify the Commission within two weeks of crossing the threshold. Luxembourg, like every member state, is subject to that regime.
In the United States, no federal law targets the development of frontier models. The TAKE IT DOWN Act, signed in May 2025 and whose removal obligations have applied to platforms since 19 May 2026, deals with non-consensual intimate content, deepfakes included, and touches neither the training nor the safety testing of large models. In December 2025 the Trump administration tasked, by executive order, a Department of Justice task force with challenging state laws it deems contrary to its policy, and on 2 June 2026 it established, by a second executive order, a voluntary framework for early government access, up to thirty days before release, to models with advanced cyber capabilities; the text rules out any mandatory licensing or pre-clearance. The bill announced on 3 September by Senator Bernie Sanders and Representative Greg Casar takes the opposite course. It would ban the development or deployment of a superintelligence, impose a temporary pause on advanced development until a new federal regulator has set safety rules, and create a cabinet-level AI supervision agency. Its chances of passing remain slim in a Republican-majority Congress. The gap between the two sides of the Atlantic is written into European law. The US Congress has not yet adopted an equivalent legal foundation.
Auditing what is already deployed remains the best available answer
Hubinger himself puts the risk from current models at a low level, consistent with the Risk Report he cites. His concern is a future superintelligence arising from recursive self-improvement, far removed from the tools a company deploys today to write, analyse or automate. Confusing the two levels leads either to ignoring a real governance debate, or to giving up, out of excessive caution, on proven uses that have nothing to do with the scenario worrying researchers.
The admission coming out of Anthropic justifies an immediate operational requirement. The same discipline the industry is trying to impose on itself for the next step should already frame what is deployed today: genuine traceability of automated decisions, and defined human oversight of uses that carry consequences. Those requirements existed before this week. Hubinger's admission makes them harder to postpone.
AIxH's view
Documenting what is running in an organisation, setting the limits of autonomy for each system and checking that governance follows usage rather than the other way round: that is the work of our AI audit and integration in Luxembourg service, which we carry out for companies in the country and internationally. The week the AI industry has just been through confirms the urgency of doing it before the next generation of systems makes the exercise harder.
