Anthropic safety lead warns of significant risk from advanced AI
Following the resignation of a colleague, a senior safety researcher at Anthropic stated there is a greater than 10% chance that artificial intelligence could pose an existential threat to humanity by the end of the decade. This warning adds to a growing discourse among experts regarding the safety and regulation of advanced AI systems.
First reported 5 days ago · latest update 5 days ago
In his post, which has been viewed more than 10 million times, Hubinger said "we really do earnestly believe" AI poses a species-ending risk to humans.
"I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," he added.
Leading figures in the AI field have been raising the alarm about the safety threat the tech poses for years, with the heads of OpenAI, Google Deepmind and Anthropic saying as much in 2023.
But those warnings have become much more stark in recent weeks, as evidence emerges that firms may be struggling to control AI.
Over the summer, there were a string of incidents where AI agents - AI systems that are allowed to operate autonomously - carried out cyber-attacks.
OpenAI, Anthropic and Meta all disclosed hacks carried out by their AI tools.
And in September, OpenAI's chief scientist Jakub Pachocki called for "extreme caution" over AI's progress, warning more intervention may be needed to ensure "humans remain in control of the future".
Major figures in the space have been calling for AI development to be slowed in recent months, including Anthropic bosses Dario Amodei and Jared Kaplan.
In an open letter signed by 1,300 staff members of AI firms, external, they called for the US government to "support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development".
Citations · 3 reports from 3 outlets
Tap a citation to read it above, right here on T.A.M.
Anthropic safety researcher says more than 10% chance AI 'could kill all humans'
It is the latest in a series of increasing warnings about the safety threat posed by artificial intelligence.
5 days agoAnthropic researcher says AI has more than 10% chance of 'killing all humans' after colleague quits
An Anthropic safety researcher said there is a greater than 10% chance AI could "kill all humans" after a former colleague quits over safety concerns.
5 days agoMore than 1 in 10 chance AI ‘could kill all humans,’ says Anthropic safety lead after colleague quits
A senior Anthropic safety researcher has said there is more than a 10 percent chance artificial intelligence "could kill all humans" by the end of the decade, just hours after a colleague resigned over fears the AI lab and its rivals are carelessly racing to build "superhuman systems" they cannot control. In a post on […]
5 days ago · Robert Hart