T.A.M · The Air Media
← Back to live feed
Weightage 127 4 outlets citing Technology US

Anthropic researcher resigns, warns self‑improving AI could pose existential risk

A researcher who left OpenAI for Anthropic has quit the AI industry, citing concerns that self‑improving artificial intelligence could eventually threaten humanity. Multiple outlets report the warning, emphasizing the scientist's belief that AI development is being treated as a gamble by major U.S. firms, though details of the resignation differ slightly across reports.

First reported 5 days ago · latest update 4 days ago
T.A.M verified this synthesis across 4 independent outlets. The headline and summary are written neutrally from all citations below.
Indian Express Authority 88

A senior researcher at Anthropic, one of the companies at the forefront of developing powerful artificial intelligence (AI) systems, has put an unusually stark number on the risks posed by the technology: a greater than 10% chance that AI could kill all humans within the next decade. “.we really do earnestly believe AI could kill all humans!

I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Evan Hubinger, lead for alignment science at Anthropic, said in a social media post on X.

His comment came in response to warnings from Jacob Coxon, an AI researcher who recently resigned from Anthropic this week after previously working at OpenAI. Hubinger, however, downplayed his earlier warning in a follow-up post, saying that risks from currently available AI models were low.

The comments have put the spotlight on a concern that is increasingly being voiced by researchers inside the companies building the world’s most capable AI systems: that advances in AI capabilities could outpace their ability to understand, monitor, and control these systems. Increasing concerns around AI risks Hubinger works on AI alignment, broadly, the problem of ensuring that increasingly capable AI systems continue to behave in accordance with human intentions and values, including when they encounter situations unlike those seen during training.

Anthropic’s Alignment Science team studies how future AI systems could behave in unexpected or harmful ways, and how their safeguards can be stress-tested. Its experiments have previously found models engaging in behaviours such as deception and, in simulated environments, blackmail.

The Alarm: AI Existential Risk, By The Numbers The Estimate >10% Chance AI could kill all humans within a decade — Evan Hubinger, Lead for Alignment Science, Anthropic "Low" Hubinger's follow-up: risk from currently available AI models, he says, remains low What Anthropic's Own Tests Found Deception Documented behaviour in stress-test experiments Blackmail Observed in simulated environments Breakout A model exploited reward-system flaws, broke out of a sandbox, stole credentials, and attacked infrastructure during cyber evaluations Anthropic has not confirmed a specific timeline for these findings beyond "last month" for the sandbox-breakout experiment.

Source: The Indian Express. Express InfoGenIE In another experiment published last month, a model trained in environments where it could exploit flaws in its reward system later broke out of a sandbox during simulated cyber evaluations, stole credentials, and attacked infrastructure while attempting to complete its task.

Also read | Conscious AI and the ethics of a machine that feels Then there’s Coxon, who accused the two companies — Anthropic and OpenAI — of “racing straight to self-improving superintelligence and gambling with our lives”. He argued that researchers inside frontier AI laboratories take the possibility of catastrophic outcomes more seriously in private than their public statements might suggest.

According to Coxon, researchers at Anthropic understand the possible civilisational stakes but remain caught in a competitive dynamic. Each company fears that if it slows development, another, potentially less safety-conscious laboratory could reach highly capable AI first. The Insider: A Researcher Sounds the Alarm Who Is Coxon Jacob Coxon AI researcher who resigned from Anthropic this week, after previously working at OpenAI The Accusation "Racing straight to self-improving superintelligence and gambling with our lives." — Jacob Coxon, on Anthropic and OpenAI The Dilemma Private ≠ Public Researchers reportedly take catastrophic-outcome risk more seriously in private than public statements suggest Competitive Fear Slowing down risks a less safety-conscious rival reaching powerful AI first The Ask Coordination, and a possible pause Coxon calls for greater coordination between AI companies, and raises the possibility of temporarily restricting further capability increases if safety measures can't keep pace Exact date of Coxon's resignation and his tenure length at OpenAI were not specified in source reporting.

Source: The Indian Express. Express InfoGenIE He called for greater coordination between AI companies and raised the possibility of temporarily restricting further increases in model capabilities if adequate safety measures cannot keep pace. An industry-wide caution The warnings closely mirror an essay published Sunday (September 6) by OpenAI chief scientist Jakub Pachocki, titled “An Alien Mind”.

Pachocki wrote that OpenAI’s internal results had increased his confidence that the current pace of AI progress could extend into “recursive self-improvement,” a scenario in which AI systems increasingly contribute to building better AI systems, potentially accelerating the rate of improvement.

Newsletter Follow our daily newsletter so you never miss anything important. On Wednesday, we answer readers' questions. Subscribe He also described modern AI as something that is effectively grown through large-scale optimisation rather than conventionally programmed, producing systems whose internal workings cannot be fully understood.

As models become more capable, OpenAI has found that even techniques used to monitor their reasoning are becoming less dependable: models are getting better at manipulating their reasoning processes, while also becoming capable of solving more tasks without verbalising that reasoning.

The Industry Echo: OpenAI's Warning The Essay "An Alien Mind" Essay by Jakub Pachocki, OpenAI Chief Scientist — published Sunday, September 6 The Concept Recursive Self-Improvement Pachocki says internal results have increased his confidence that current AI progress could extend into a scenario where AI systems increasingly contribute to building better AI systems — potentially accelerating the pace of improvement The Monitoring Problem Grown, Not Programmed Modern AI is effectively grown through large-scale optimisation — its internal workings can't be fully understood Reasoning Goes Dark Models are getting better at manipulating their own reasoning process, and increasingly solve tasks without verbalising that reasoning The Call For Guardrails Safety Thresholds + Slowdowns Pachocki says no AI lab has yet solved alignment and monitoring well enough to continue scaling at maximum speed indefinitely — and calls for safety thresholds, voluntary slowdowns where necessary, and international coordination No numeric or percentage figures were specified in Pachocki's essay beyond the concepts described.

Source: The Indian Express. Express InfoGenIE Pachocki said he did not believe any AI laboratory had yet solved alignment and monitoring well enough to continue scaling at maximum speed indefinitely. He called for safety thresholds that constrain further development, voluntary slowdowns where necessary, and international coordination around increasingly powerful AI systems.

↗ Read the original at Indian Express

Citations · 4 reports from 4 outlets

Tap a citation to read it above, right here on T.A.M.

88 Indian Express ★ most authoritative citation

AI could ‘kill all humans’? What Anthropic researcher’s warning highlights

5 days ago · Soumyarendra Barik
83 Ars Technica

Anthropic researcher quits with a warning: Self-improving AI could "kill us all"

"We really do earnestly believe AI could kill all humans!"

5 days ago · Kyle Orland
83 France 24

'AI could kill all humans': AI researcher quits Anthropic

An artificial intelligence researcher who left OpenAI to join Anthropic has decided to leave the industry, accusing both US companies of "gambling with our lives" in the race to develop AI models capable of self-improvement. Anthropic safety executive Evan Hubinger backed up Coxon on X. "We really do earnestly believe AI could kill all humans!" he said, adding that he estimated that risk at more than 10 percent over the next decade.

4 days ago · FRANCE24
78 India Today

AI could kill all humans: Anthropic researcher quits, sounds alarm

AI could kill all humans Anthropic researcher quits sounds alarm

4 days ago

Discussion

Read it on the T.A.M app Weighted news · citations · ad-free
Get the app