INQUIRING LINE

Google searches that start with a full question get fewer clicks anyway — so is AI really to blame for the drop?

Could longer question-based searches naturally have lower click rates anyway?

This explores whether the drop in clicks seen when Google shows AI summaries might be partly a selection effect: longer, question-style searches may get fewer clicks whether or not a summary appears, because of the kind of question being asked.


This explores whether lower click rates on AI-summary searches might partly come from which searches get summaries, rather than only from the summaries themselves. The headline number is stark: in Pew's browsing data, people clicked a result link 8% of the time when an AI summary appeared, against 15% without one, and those sessions were more likely to end with no click at all Do AI summaries on Google reduce clicks to actual websites?. Your question asks the right thing. If summaries mostly appear on long, question-shaped searches, and those searches would have drawn fewer clicks anyway, then part of the gap belongs to the type of query, not to the AI. The corpus doesn't contain a study that separates these two effects, so it can't settle the question directly. It does give good reasons to think query shape matters.

The strongest lateral evidence comes from retrieval research, which treats the form of a question as a signal in its own right. One line of work finds that non-factoid questions (comparisons, debates, 'why' and experience questions) need different retrieval strategies from simple lookups. Some need to be broken into parts, and some need evidence gathered from several angles Does question type determine the right retrieval strategy?. A question like that rarely has one page that answers it, so the classic 'click the top link' behavior fits it poorly from the start. A related finding: lightweight surface features of a question, such as its length and structure, predict whether retrieval is needed about as well as expensive model-uncertainty methods Can question features alone predict when to retrieve?. If a question's shape tells a system how much searching it needs, it plausibly tells you something about how a person will search, too.

Research on deep research agents points the same way from the other side. Answer quality on complex questions keeps improving with more search steps, with diminishing returns, much as it does with more reasoning Does search budget scale like reasoning tokens for answer quality? Do search steps follow the same scaling rules as reasoning tokens?. Hard questions are answered across many sources rather than by one page. For a human, that could mean reformulating the query, giving up, or settling for a partial answer. None of those registers as a click on that particular search, so lower click rates for long questions could reflect people moving between searches rather than being satisfied.

One more twist cuts against treating a 'no click' as a sign the answer was good. In AI search arenas, users rated responses with more citations as more trustworthy, even when the extra citations were irrelevant Do users trust citations more when there are simply more of them?. Citations can work as a reassurance signal rather than an invitation to click. So even after separating out query type, an AI summary might cut clicks by making people feel informed instead of by actually informing them. The real test would compare clicks on matched queries of the same length and question type, with and without a summary. The corpus doesn't have that comparison, and anyone using the Pew figure as a pure causal effect should be asked for it.


Sources 6 notes

Do AI summaries on Google reduce clicks to actual websites?

Pew's analysis of 68,879 Google searches found users clicked search result links 8% of the time when an AI summary appeared, versus 15% without one. Sessions were also 10 percentage points more likely to end without any clicks.

Does question type determine the right retrieval strategy?

Research shows non-factoid questions split into five types, each requiring different retrieval and aggregation methods. Evidence-based questions suit standard RAG, while debate and comparison need aspect-specific retrieval, and experience/reason questions need decomposition or filtering strategies.

Can question features alone predict when to retrieve?

Learned predictors using 27 lightweight external question features match complex uncertainty-based methods on overall performance while costing far less, and outperform them on complex questions across 6 QA datasets.

Does search budget scale like reasoning tokens for answer quality?

Agentic deep research shows monotonic-to-diminishing-returns curves for search iterations, matching reasoning token scaling. This creates a new inference-compute axis: models can trade off reasoning budget against search budget to optimize answer quality.

Do search steps follow the same scaling rules as reasoning tokens?

Deep research agents improve with more search steps in a pattern mirroring the reasoning-token relationship, with both exhibiting diminishing returns. This reveals a new inference-compute axis beyond model capability alone.

Show all 6 sources
Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.