METR’s first Frontier Risk Report
Anthropic, Google, Meta, and OpenAI let METR test internal models with chain-of-thought access and review non-public evidence about agent control risks
Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test their best internal models with CoT access, (2) review non-public info about capabilities, alignment, and control. The res