Which jobs could be done by AI on paper, yet barely use it in real life?
Which occupations show the sharpest gap between AI capability and actual adoption?
This explores which kinds of jobs look most AI-capable on paper but see the least real AI use. The corpus has no ranked list of occupations, so this reads the evidence for where the gap likely sits.
This explores which kinds of jobs look most AI-capable on paper but see the least real AI use. The corpus doesn't rank occupations by that gap, but it shows where the gap is smallest, and that narrows the search. Delegated AI work, where people have handed tasks to AI inside structured workflows, concentrates in information-intensive occupations. It follows what the technology can do rather than how popular chatbots are Where have workers actually delegated tasks to AI?. So capability and adoption line up best in information work, and the sharpest gaps are probably elsewhere.
The first place to look is jobs where the capability was never real, only benchmarked. An analysis of 960 real occupational workflows found agents winning abstract contests but failing long-horizon professional tasks. The authors' point is that the field measured contests rather than work Why do agent benchmarks not predict real economic value?. In a simulated company, leading agents finish only about 30% of tasks. Their failures cluster in social interaction, navigating professional software, and domain-specific knowledge Why do AI agents fail at workplace social interaction?. The occupations built on those three things, such as client-facing coordination and work inside specialized professional tools, are where headline capability and usable capability diverge most. That is my inference from the failure modes, not a finding the notes make about specific jobs.
Even where capability exists, other things hold adoption back. A historical analysis finds capable agents stall without five ecosystem conditions: value generation, personalization, trustworthiness, social acceptability, and standardization Why do capable AI agents still fail in real deployments?. Trust matters most in professions whose authority comes from peer validation. Expertise is confirmed through a track record within a community of experts, and the corpus argues AI can't join that circle however accurate it is Can AI ever gain expert community trust through participation?. For those fields the gap may be about who is allowed to rely on the output, not about what the model can do.
Exposure is not adoption, though, and the corpus suggests the gap falls unevenly. AI-skill hiring clusters in STEM around Python, SQL, and machine learning, while non-technical occupations drift away from that core Is AI creating common skills across jobs or deepening divisions?. In female-dominated occupations, exposure spreads across all skill and wage levels. In male-dominated ones it concentrates among high-paid workers. That leaves lower-paid workers exposed with fewer resources to adapt Does AI exposure hit low-wage workers harder in some fields?. Firms differ too: highly exposed firms replace freelance labor with AI faster and more cheaply, which points to returns to scale rather than an even spread of the technology Do firms substitute labor for AI at different rates?. The best-supported reading is that the widest gaps sit in non-technical, socially heavy, long-horizon professional work. No note measures that directly, so it's an open question in this library.
Sources 8 notes
Workers have committed AI tasks to structured workflows primarily in information-intensive occupations, following technical capability more than conversational LLM adoption. This gradient differs sharply from routine-task automation predictions and wage patterns reverse at advanced degree levels.
ALE's analysis of 960 real occupational workflows shows agents excel at abstract contests but fail long-horizon professional tasks. The gap is not model capability but benchmark design—the field optimizes what it measures, and it has measured contests rather than work.
TheAgentCompany benchmark shows leading agents achieve 30% task completion in a simulated workplace. Social interaction, professional UI navigation, and domain-specific knowledge are the three primary failure modes, with multi-turn task performance consistently dropping to 35% across enterprise settings.
Historical analysis from GPS to modern AI shows agent failures consistently result from absent ecosystem conditions—value generation, personalization, trustworthiness, social acceptability, and standardization—rather than capability gaps. Even highly capable systems stall without these five conditions.
Expertise is validated through social participation and track record within expert communities, not individual accuracy alone. AI cannot enter this validation circle because it lacks social embeddedness, testable judgment history, and ability to participate in the consensus-building processes that define expert paradigms.
Show all 8 sources
Vacancy data from ten countries show AI skill demand concentrating heavily within STEM occupations around Python, SQL, machine learning, and data analysis, while non-technical occupations diverge from this core rather than converge toward it.
AI exposure concentrates among high-skilled, high-paid workers in male-dominated occupations but spreads evenly across all skill levels in female-dominated ones. This means lower-paid, lower-skilled women face disproportionate exposure despite having fewer resources to adapt.
Higher AI-exposed firms replace online labor marketplace workers with AI tools faster and at lower cost than less-exposed firms, suggesting returns to scale in internal AI capability rather than uniform technology diffusion.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- When AI Enters the Workplace, Who Faces Greater Risks? A Gendered Analysis
- Who Delegates to AI? Evidence from Agent Configurations in Github
- Artificial Intelligence and the Labor Market∗
- Occupational Convergence or Divergence? Mapping Labor Market Structural Shifts Driven by AI Penetration
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Gdpval: Evaluating Ai Model Performance On Real-world Economically Valuable Tasks
- Future of Work with AI Agents: Auditing Automation and Augmentation Potential across the U.S. Workforce
- TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks