AI, networks and Mechanical Turks
Source: Benedict Evans, Benedict Evans (newsletter/essays) · 2025-11-23
The limitation of this, though, is that none of these systems really know why you watched those, or bought those, or looked at those, and they don't really know what those things are. Amazon has close to a billion SKUs, but all it knows about them is the metadata typed in by humans and some level of purchase correlation. Instagram or TikTok have the same problem. These systems have correlation across SKUs, and down the funnel, but they don’t know why, just as a dog knows that the sound of door keys has high correlation with a walk, but doesn’t know what keys are.
But an LLM is, at a minimum, a step change in automated understanding of both what and why. The model can look at those words, images, videos and products, and all that metadata, and connect them to patterns that have some kind of understanding, or at any rate, some vastly broader kind of correlation.
So, YouTube can say “you watch car chase videos. Here’s a video that seems to have a car chase”. But this can also bring new kinds of correlations (or, if you prefer, ‘understanding’). Today, Amazon will know that if you buy packing tape, you might want bubble wrap. It should know that if you buy those, you might also want lightbulbs and maybe smoke detectors, because you’re moving house. But an LLM might know to show you an ad for home insurance and broadband, things you probably couldn’t infer from Amazon’s purchasing data.
You don't necessarily need all of that user base to do this, either - or rather, it doesn’t need to be your user base, because you might not need to build your own Mechanical Turk. If this kind of knowledge generalises enough, it might just be an API call from a world model. You can rent the cold start. Where Amazon or TikTok built recommendation systems by looking at what people do on Amazon and TikTok, that might just be the inference of a general-purpose LLM that anyone can plug into their product. You move the ‘human in the loop’: now the human, and the mechanical turk, is in the creation of all that training data over the past few hundred years. You shift the point of leverage.
Generalising that problem, Google, Amazon, Meta, TikTok, Tinder and a bunch of other companies (Uber, Doordash, Venmo) all know something about you. But they each have a very partial view, like the story of the blind men feeling an elephant. Your phone has a much wider view, in theory, although Apple and Google are very cautious in what they do with that (Chinese Android OEMs are much more aggressive here), but it’s limited in different ways: your phone could know what you bought on Amazon or what images you saw in Instagram, but it wouldn’t see the graph that led Meta to show them to you. Now an agentic LLM assistant might become another blind man feeling the elephant, in different ways: it knows different things, but as you use it it also knows different thing about you, drawing new conclusions, especially if it’s buying things on Amazon and Instacart for you, and you’re using its browser, and its wearable.
But I do think that this is a new turn on a problem I’ve written a lot about in the past: the internet removed all of the old filters, curation and editing, so that now we have effectively infinite product, infinite media and infinite retail, and no way to or find or see what we don’t know. The filters we had from the internet were very imperfect, and now we have a radically new and different kind of filter. That seems like a bigger question than replacing Google.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can AI systems reliably guide voters without introducing political bias?- Is ChatGPT adoption concentrated among already-advantaged, highly-paid workers?
- Does ChatGPT displace search engines or question-and-answer platforms?
- Why does writing dominate work-related ChatGPT use compared to other tasks?
- What earnings or employment changes follow ChatGPT adoption in real datasets?
- Why did Upwork freelancers lose earnings after ChatGPT's release?
- How much referral traffic do ChatGPT and Google each send to publishers?
- How do ingroup disagreement and outgroup agreement differ in their depolarizing mechanisms?
- Does designing chatbots to satisfy partisan users prevent them from reducing polarization?
- Do chatbots actually give consistent voting recommendations regardless of user input?
- Can chatbots that offer open-ended advice avoid the same polarizing collapse?
- Should election regulators require chatbots to refuse voting advice entirely?
- Why do chatbots trained on internet data show consistent political bias?
- Do AI systems mirror user political views when they choose to engage?
- How did imperfect AI representations still correct partisan misperceptions effectively?
- Are other Google SERP features like Featured Snippets also reducing organic clicks?
- Can researchers isolate AI Overviews from other confounds affecting CTR trends?
- What happens to publisher revenue when search referral traffic collapses?
- Do smaller ad-dependent publishers face larger traffic losses from AI summaries?
- Which search queries trigger AI summaries most often on Google?
- Do users click links within AI summaries or end sessions instead?
- Do searchers prefer clarity about AI involvement when viewing search overviews?
- How much do AI Overviews currently appear in Google search results?