SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Does Anthropic's survey adequately support its automated R&D risk conclusion?

Anthropic's Risk Report surveyed model capabilities in automated R&D and concluded risks are very low. But does the survey design and execution actually justify that strong conclusion, or does it have flaws that undermine it?

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

METR's review of the "Risks from automated R&D" section in Anthropic's February 2026 Risk Report finds the report's case too weak to carry its own bottom line. The report "makes the argument that catastrophic risk from Claude Opus 4.6 or a less capable Anthropic model automating R&D in any domain is very low." METR's verdict on the argument itself: "We do not think the report adequately supports its conclusion." It flags "analytical rigor" problems in the report's model-use survey — "issues including sample size, question granularity, survey framing" — and says the overall argument "misses the possibility of substantial AI R&D acceleration before its full automation." On "adequacy of information," METR finds Anthropic "summarizes the survey results in a way that miscounts one missing response to a question as a negative response."

The mechanism of the review is a split between evidence and conclusion. "If we had to solely rely on the evidence presented by Anthropic in the original Risk Report, we would likely disagree with the report's conclusion." What changes METR's mind is evidence from outside that survey: "additional evidence indicating that the model is incapable of R&D in key domains, including the results of METR evaluations and the lack of public reports of the model automating any key domain." On that separate basis, "we agree with the bottom-line conclusion of the report... but we think the evidence presented in the report is inadequate to establish this." METR also recommends Anthropic widen the internal survey — bigger sample, more granular response options, different framing — and report other, more leading indicators of AI progress.

The report METR is reviewing sits next to two self-reported Anthropic measures of the same automation trend. Is AI development already being handed to AI systems? cites task-length growth, code-authorship share and a speedup test as evidence of delegation to AI. How fast is AI accelerating its own development inside labs? goes further and names the fix directly: that snapshot, "by the source's own account, would need third-party verification... before it could show how the pace is changing." METR's review is that outside check, applied to the companion risk-assessment survey, and it reaches a sharper verdict than either note anticipates: not that the self-reported numbers are wrong, but that a survey built on them cannot establish a "very low risk" conclusion, regardless of whether that conclusion happens to be correct.

The excerpt is an executive summary: it names the problem categories — sample size, question granularity, survey framing, prior METR research on calibration difficulty, the miscounted missing response — without giving the underlying numbers, question text, or response counts, which sit in the full PDF reviews it points to but doesn't include. It also doesn't say how the miscounted response changed the report's published figures, or what a corrected survey would show. The narrower result that survives is this: a lab's self-administered survey of its own models, even when it points toward a conclusion that later holds up on independent evaluation, is not itself sufficient evidence for a catastrophic-risk judgment.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Should models ask for clarification when facing ambiguous or under-specified information?

Related concepts in this collection 2

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 110 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

METR agrees Anthropic's low-risk conclusion on automated R&D but finds the survey evidence supporting it inadequate