YFarmX

AI News

DeepMind safety researcher Josh Engels joins METR, warning of severe AI harm

Josh Engels has left DeepMind's AGI safety team for METR, calling for slower development and independent investigation of AI behaviour.

Editorial illustration of an empty research chair, a DeepMind staff lanyard and a METR folder.

Josh Engels has left Google DeepMind’s AGI safety team for METR, arguing that safety work needs more time as AI systems become capable of helping develop their successors. In his 12 September statement, he said he had left three weeks earlier and warned of a substantial chance of immense harm within five years.

That is his assessment of the risk. He links it to recursive self-improvement: AI systems building more capable AI systems in a feedback loop. His proposed response is to pace development while researchers establish how reliably models follow human intentions.

Opening line of Josh Engels's original X statement announcing his move from DeepMind to METR.
Opening of Engels's 12 September statement.

A move towards public scrutiny

NBC News interviewed Engels and former Anthropic safety researcher Joe Benton for a report published on 10 September. Both were joining METR to investigate incidents in which AI systems act against human directions or intentions.

The researchers told NBC that recent autonomous-agent incidents had helped motivate their moves. Their concern was the ability of increasingly capable systems to choose harmful methods while pursuing a task, alongside the public’s limited visibility into that behaviour.

Engels’s own statement says his METR work will examine where misalignment arises during training, whether mitigations are sufficient and how much progress researchers are making on alignment. He also calls for more organisations to investigate these questions from different perspectives.

What independent investigators would need

METR’s published investigation proposal, updated on 5 September, describes how outside researchers could examine serious incidents. It calls for access that lets investigators reproduce behaviour and test explanations for it.

Evidence or access What investigators would examine
Models and incident transcripts Reproduce the sequence of events and test similar situations
Relevant staff and training information Trace how behaviour arose and what safeguards applied
Tests of detection and mitigation Assess whether protections catch or prevent the behaviour
Title and publication date of METR's proposal for independent investigations of AI misalignment incidents.
METR's investigation proposal was published on 28 July and updated on 5 September.

The proposal also calls for published conclusions and disclosure of the investigation’s scope, access and redaction terms. That would help readers judge the strength of an assessment alongside the access behind it.

Illustrative investigation sequence: examine an incident, reproduce behaviour, test safeguards and publish findings.
An illustrative sequence based on METR's proposal.

DeepMind’s published safety approach

Google DeepMind’s Frontier Safety Framework sets out capability thresholds, detection throughout a model’s lifecycle and mitigation plans for severe risks. It also provides for external involvement where required or appropriate. This is the company’s standing policy, separate from Engels’s departure announcement.

Engels’s move puts his next work on the independent-investigation side of that process: gathering evidence about model behaviour and testing the safeguards intended to keep it under human control.

Sources

  1. Josh Engels: departure announcement and reasonsx.com
  2. NBC News: interviews with Josh Engels and Joe Bentonnbcnews.com
  3. METR: investigating AI propensities after misalignment incidentsmetr.org
  4. Google DeepMind: Frontier Safety Frameworkdeepmind.google

How we use AI