Can expert-written questions resist web-assisted non-expert answering?
GPQA tested whether graduate-level multiple-choice questions written by domain experts remain difficult even when non-experts have unrestricted internet access. This matters for building benchmarks that can supervise AI systems on tasks beyond typical human reach.
The GPQA paper reports 448 multiple-choice questions in biology, physics and chemistry, written by domain experts and checked so that they are hard for people outside the field. Experts "who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect)". "Highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web." The strongest GPT-4-based baseline reaches 39%. The figures are the paper's own measurements, taken from its contractor pool; the excerpt names no outside source for them.
The construction explains the gap. Each question went through question writing, a first expert validation, revision by the writer, a second expert validation, and then three non-expert validators who tried to answer it with internet access. The paper's stated goal is scalable oversight: as models answer harder questions, annotators "may be difficult" to supervise even when they are "themselves skilled and knowledgeable". The authors want testbeds "as close as possible to the edge of human expertise", and they claim GPQA meets seven of Irving and Askell's nine desiderata for scalable-oversight datasets.
Against the nearest notes, the GPQA result is the human baseline that oversight work would need. Why does assisted accuracy capture only half the LLM gain? measures what happens when a person works with a model, and finds assisted accuracy below the better component. GPQA measures no assisted condition, only unaided experts and web-equipped non-experts, so the two studies test different things and should not be merged. The debate note's condition, that When does debate actually improve reasoning accuracy?, is useful here: GPQA's non-experts can search, yet still reach 34%, which suggests that retrieval alone does not give a non-expert the checking that debate needs. That is an inference from the numbers, not something the paper tests. It also differs from Can crowdsourced votes reliably rank language models?, where crowd votes track expert raters on preference. GPQA asks about correctness at graduate level, and there skilled non-experts diverge from experts.
The excerpt does not establish that GPQA works as an oversight testbed. It reports no oversight experiment, so the paper's hope that the benchmark "can help devise ways for human experts to reliably get truthful information from AI systems" is a proposal, not a result. The non-expert figure is also a ceiling, not a typical rate: the authors say "our numbers will likely not directly translate to realistic settings". The discussion adds that non-expert accuracy is biased downward, because it was used to select the subsets. The set is small at 448 items, which the limitations section says gives little power for accuracy comparisons unless the effect is large, and the experts were sourced through Upwork with no enforced demographic mix. The authors "make no claim" that the set is representative.
At the strength the evidence allows, the result supports one claim: these questions are hard for both skilled non-experts with web access and a frontier model, so GPQA can serve as a hard accuracy benchmark near the edge of expert knowledge. Whether it can validate any oversight method is still open.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How effectively can test-time voting aggregate diverse reasoning samples? How do educators verify student capability when AI can produce indistinguishable work?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why does assisted accuracy capture only half the LLM gain?
When an AI system improves on a task, how much of that improvement actually reaches people using it with the AI? Understanding this gap matters because it shows whether complementary strengths automatically translate to better team performance.
the human-with-model setting that GPQA's oversight framing points at; GPQA itself measures only unaided conditions
-
When does debate actually improve reasoning accuracy?
Multi-agent debate shows promise for reasoning tasks, but under what conditions does it help versus hurt? The research explores whether debate amplifies errors when evidence verification is missing.
GPQA's non-experts can search yet reach 34%, which bears on whether retrieval alone supplies the evidence checking debate needs
-
Can crowdsourced votes reliably rank language models?
Explores whether large-scale human preference voting from casual users produces valid model rankings comparable to expert judgment, and what makes such crowdsourced evaluation trustworthy at scale.
contrast: crowds agree with experts on preference, while skilled non-experts diverge from experts on graduate-level correctness
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Prompting Science Report 4: Playing Pretend: Expert Personas Don't Improve Factual Accuracy
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
- OpenThoughts: Data Recipes for Reasoning Models
- Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
- Evaluating Large Language Models in Theory of Mind Tasks
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- Evaluating the psychometric properties of ChatGPT-generated questions
Original note title
GPQA's expert-written questions stay Google-proof for skilled non-experts — PhDs reach 65% and non-expert validators 34%, a scalable-oversight testbed