Did superhuman AI actually improve Go players' decision quality?
After AlphaGo's breakthrough, professional Go players made better moves and tried more novel strategies. But did AI exposure directly cause this improvement, or did players simply memorize AI moves?
Decision quality among professional Go players rose after superhuman AI arrived, and the rise came with more novel play. The authors date the "advent of superhuman AI" to a series of events between 2016 and 2017, anchored by AlphaGo's March 15, 2016 defeat of a human world champion. Across more than 5.8 million move decisions from 1950 to 2021, they score each actual move with KataGo, a superhuman program. It simulates 10,000 game patterns per decision and compares the win rate of the human move with that of the counterfactual AI move, producing a Decision Quality Index. The excerpt reports that "humans began to make significantly better decisions following the advent of superhuman AI." Novelty, measured as each game's first historically novel move, also rose: novel decisions "occurred more frequently and became associated with higher decision quality after the advent of superhuman AI."
The mechanism the authors propose is that superhuman AI "can ultimately unearth superior solutions previously neglected by human decision-makers who may be focused on familiar solutions." Because AI can identify optimal decisions "free of human biases (especially when it is trained via self-play)," its moves widen the options players consider, and the paper suggests they "induced them to explore novel moves." The main alternative explanation is memorization of AI moves. The authors answer it by excluding every human move that matched the optimal AI decision, and a "sharp increase in decision quality" remains. They stop short of making novelty the whole story: the improvement "may be partly explained by increased novelty."
The library's nearest notes frame human-AI interaction differently. In the Atria Dawn analysis, agents propose methods while humans keep most final decisions and steer exploration, so exploratory judgment stays with people. The Go paper shows a different locus: the human side of exploration itself widened after a machine showed unfamiliar moves. The hybrid-society study (Do humans learn to prefer AI partners over time?) also shows humans adapting to AI through repeated interaction, but the adaptation there is partner choice, where here it is move choice and its measured quality. The long-horizon agent study (Do frontier AI agents actually conduct novel research or just optimize?) finds genuine novelty rare in agent output. Read together, the cases suggest novelty has shown up in human play after AI exposure but not in agents' own work. That is a contrast of scope, since the Go excerpt does not test agent novelty.
The excerpt establishes an association, not the full causal chain. The authors list as open whether AI raised novelty and thereby raised quality, and through which other mechanism it might have acted. The quality measure is also a machine estimate: moves are judged better by KataGo's win rates, not by a proof assistant, expert review or players' own explanations. The excerpt does not say that players understood why the new moves were better. It can show that Go decisions got better and that novelty tracked the change, but not that the people making those decisions grasped the reasons. At the strength this evidence allows, "better" here means better by a superhuman evaluator, and whether the improvement was understood is a separate question the excerpt leaves open.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do AI systems determine and balance multiple competing objectives? How can evaluations be made robust against model reward hacking?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How should AI agents and humans divide research tasks?
In building its own foundation model, Atria Dawn studied how to split work between agents and human researchers. Understanding this division matters for designing effective human-AI collaboration in technical R&D.
contrasts: there humans steer exploration, while here human exploration itself widened after AI moves appeared.
-
Do humans learn to prefer AI partners over time?
Exploring whether repeated interaction with AI agents shifts human partner selection despite initial bias against machines. This matters because it tests whether behavioral performance can overcome identity-based resistance in hybrid societies.
extends: both show humans adapting to AI through repeated exposure, here in move choice and decision quality rather than partner choice.
-
Do frontier AI agents actually conduct novel research or just optimize?
Exploring whether current long-horizon research agents generate genuine methodological novelty or primarily recombine established techniques. This matters for understanding how close we are to recursive self-improvement through AI.
contrasts: that study finds agent novelty rare, while here novelty surfaces in human play after AI exposure.
-
Can humans learn chess concepts that AlphaZero discovered alone?
Do grandmasters improve on new puzzles after seeing AlphaZero's solutions to similar positions? The question tests whether superhuman chess knowledge can transfer from machine to human player.
Evidence for: four grandmasters improved on puzzles built from AlphaZero's concept vectors, suggesting AI-derived novelty can lift human decisions, in chess rather than Go
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Superhuman Artificial Intelligence Can Improve Human Decision Making by Increasing Novelty
- Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero
- Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
- Teaching Large Language Models to Reason with Reinforcement Learning
- How Well Can AI Do Strategy? Empirical Benchmarking Using Strategy Simulations
- A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1
- Tree Search for Language Model Agents
- Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
Original note title
decision quality among professional Go players rose after superhuman AI arrived, alongside more novel moves — novelty may partly explain the gain