SYNTHESIS NOTE
Topics›AI at Work›this note

Does AI really compress all layers of knowledge work equally?

Explores whether AI's productivity gains affect all stages of knowledge work uniformly, or whether some tasks like planning and accountability grow harder as execution becomes easier.

Synthesis note · 2026-10-09 · sourced from AI at Work

Arvind Narayanan, in an ICML keynote transcript co-framed with Sayash Kapoor's "AI as Normal Technology" work, splits software engineering into three layers: "the decide layer: understanding customer requirements, developing the specification, planning, etc," "the execute layer: the actual coding and debugging," and "the deliver layer: understanding your code deeply enough to be accountable for what you release." His claim is that only the middle layer compresses under AI, and that it "was only maybe one-third of the work to begin with." The decide and deliver layers are "not getting compressed," and he argues "the first and third layers are arguably expanding as AI compresses the middle layer."

The reasoning rests on several supporting observations in the transcript. Capability and reliability diverged in benchmarks of three frontier companies over roughly 24 months: accuracy "shot up dramatically" while reliability "only increased by five or ten percentage points" — a gap he uses to explain why deployment lags behind what capability alone would predict. He also rejects the inference that productivity gains must shrink headcount, invoking "Jevons' paradox" and the "lump-of-labor fallacy": in cases his team examined, layoffs attributed to AI occurred at companies already "under financial pressure," making AI a convenient scapegoat rather than the cause. Translation is his example of a task where AI reached near-human parity "nearly a decade ago" yet human-translator employment "has remained more or less stable," because there is "no ceiling" to translatable volume. Lawyers see more work, not less, because AI "made it a lot easier to file lawsuits." The upshot he draws is occupational, not just task-level: "programming jobs," narrowly defined around coding and debugging, have diverged in demand from "software engineering jobs," which carry decide-execute-deliver responsibility in full, and he predicts this split will recur across fields as effort shifts "from building systems to evaluating systems."

This sits closely with What makes accountable judgment scarce when AI cognition is cheap?, which names the same residual as "accountable judgment" and treats occupations as "governance bundles" rather than task lists — Narayanan's decide and deliver layers are a specific, named instance of that bundle. How should AI agents and humans divide research tasks? reports an empirical version of the same split inside one R&D project, with agents taking the execute-like role and humans the decide-like one, though at the scale of a single company rather than an occupation-wide claim. Does personal preference shape how engineers use AI tools? complicates the picture by showing that who actually keeps the decide layer in practice is often set by employer policy rather than emerging naturally from the task structure Narayanan describes.

The excerpt is a lightly edited keynote transcript, not a paper: the capability-reliability benchmark, the layoffs analysis, and the programming-versus-software-engineering divergence are each referenced as findings from the author's other work ("We looked at this in a follow-up essay," "We've written a paper") but not reproduced here with method or data, so none of the specific figures can be checked from this excerpt alone. The three-layer framework itself is presented as the author's own model for software engineering and extrapolated to other professions by prediction ("I predict that this will happen in more and more fields over time"), not by evidence from those other fields. What follows at the strength the excerpt supports is a framework worth testing occupation by occupation, not a general finding that decide and deliver layers are safe from compression everywhere.

Inquiring lines that read this note 30

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can AI research automation sustain progress through accelerating feedback loops? Does AI-assisted work increase total productivity or just shift time? How should human-AI contributions be measured, disclosed, and verified? How do writers navigate authorship and delegation with AI? How can humans maintain effective oversight as AI systems scale? Does AI deployment reduce or exacerbate workplace inequality and income instability? How does AI adoption reshape collaboration patterns in knowledge work? How do AI-exposed occupations change in employment, wages, and skills? How should humans and AI agents share control and decision-making? Does AI assistance help or harm professional skill development?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 126 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Narayanan and Kapoor argue AI compresses the execute layer of knowledge work while the decide and deliver layers expand