SYNTHESIS NOTE
Topics›AI at Work›this note

Why do most enterprise AI pilots fail to deliver returns?

MIT NANDA investigated why 95% of enterprise generative AI pilots produce no measurable profit impact. The research explores whether failure stems from weak models, regulation, or how organizations actually deploy and use these tools.

Synthesis note · 2026-10-09 · sourced from AI at Work

MIT NANDA's "GenAI Divide" report argues that enterprise generative AI investment is not failing because the underlying models are weak; it is failing because of how organizations buy and use the tools. Drawing on "52 structured interviews across enterprise stakeholders, systematic analysis of 300+ public AI initiatives and announcements, and surveys with 153 leaders," the report finds that despite "$30–40 billion in enterprise investment into GenAI," "95% of organizations are getting zero return," while "just 5% of integrated AI pilots are extracting millions in value." The split runs across both buyers (enterprise, mid-market, SMB) and builders (startups, vendors, consultancies), and the report states plainly that "this divide does not seem to be driven by model quality or regulation, but seems to be determined by approach." General-purpose chatbots such as ChatGPT and Copilot are "widely adopted" — over 80 percent of organizations have piloted them and nearly 40 percent report deployment — but the report says these tools "primarily enhance individual productivity, not P&L performance." Custom or vendor-built enterprise systems fare worse at the production stage: 60 percent of organizations evaluated them, 20 percent reached pilot, and "just 5 percent reached production."

The report locates the mechanism in a single capability gap: "most GenAI systems do not retain feedback, adapt to context, or improve over time." Static tools that need constant re-prompting stall once a workflow needs more than a scripted answer, which is why the report calls this the "pilot-to-production chasm" — generic chatbots show an "~83%" pilot-to-implementation rate because they are easy to try, but "fail in critical workflows due to lack of memory and customization." The report also credits a build/partnership effect: "external partnerships see twice the success rate of internal builds," and the organizations that cross the divide "buy rather than build, empower line managers rather than central labs, and select tools that integrate deeply while adapting over time." A separate finding the report calls the "shadow AI economy" supplies corroborating evidence for the approach argument from the opposite direction: although "only 40% of companies say they purchased an official LLM subscription, workers from over 90% of the companies we surveyed reported regular use of personal AI tools for work tasks," which the report reads as proof that "individuals can successfully cross the GenAI Divide when given access to flexible, responsive tools" even where official deployment has stalled.

That shadow-economy observation lines up with Does generative AI shift knowledge workers away from communication?, whose M365 trace data shows individual, informal usage changing what heavy users do well before any official organizational rollout shows up in the numbers — both sources find the organizational unit of adoption lagging the individual unit. The report's "pilot-to-production chasm" also matches the texture How are national lab staff actually using generative AI? documents inside one science organization: survey and interview evidence there shows use "largely experimental," sorted into copilot and early-stage workflow-agent modalities, which is the kind of case NANDA's 300-initiative count renders as an aggregate statistic. Set against those two matches, Can generative AI replace the benefits of having a human teammate? looks like the pre-registered, outcome-measured success the GenAI Divide report treats as rare — a single field experiment with a measured business-relevant result, the kind of case the report would count toward its 5%, not its 95%.

The report's own limitations section concedes its figures "are directionally accurate based on individual interviews rather than official company reporting," that "build vs. buy percentages" rest on interview responses "rather than comprehensive market data," and that its "six-month observation period may be insufficient to fully assess 'successful deployment'" for complex systems — so the headline 95%/5% split is a snapshot read from self-selected interviewees and public announcements, not an audited outcome measure, and slower-maturing deployments could shift the ratio. The report does not test whether model quality or regulation plays no role so much as rule them out by interview impression, without a controlled comparison of firms matched on everything but buy-vs-build choice. At the strength the evidence allows, the finding supports treating workflow fit and tool adaptability as more useful predictors than model capability when forecasting which GenAI investments will show returns, while the specific 95%/5% ratio should be read as this report's own directional estimate rather than an industry constant.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do real-world evaluations reveal AI capabilities that benchmarks hide?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 81 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

MIT NANDA finds 95% of enterprise generative AI pilots return no measurable P&L impact — the divide tracks approach, not model quality or regulation