The New Shape of Search: How Conversational AI Recomposes Information Seeking
The familiar search journey begins with a query and moves outward into documents, and conversational AI is commonly imagined at its mouth: ask first, then click out. Linking captured prompts and responses to the same panelists’ observed searches and pageviews, and reconstructing inactivity-defined cross-surface temporal sessions (standalone assistant surfaces; search-embedded AI such as AI Overviews and AI Mode is out of scope, since it co-occurs with the results page), we find the observed journeys more often run the other way. Content usually follows search but more often precedes assistant use. Within the same panelist, the paired difference-ofdirections between the two anchors is +20.6 [19.9, 21.3] percentage points; it persists within every coarse destination-domain stratum we can observe (semantic task and task-stage matching remain unresolved), and every headline result replicates in a second, adjacent month. Search tends to anchor the front of the observed journey; assistants sit deeper inside it. Assistant sessions are also far more often self-contained. Userweighted, 34.1% [33.5, 34.7] of assistant-containing sessions show no observed external web step (AI-first 10.5% [10.2, 10.9], AI-last 18.3% [17.8, 18.7], bridge/interleaved 37.1% [36.5, 37.7]), against 19.5% [19.2, 19.8] contained for search-centered sessions of the same users, a within-user contrast of +13.0 [12.5, 13.6] percentage points. We call this difference recomposition: activity is distributed differently across dialogue, search, and browsing, without implying that assistant use caused the difference. Assistant-contained also does not mean resolved: timestamps alone cannot establish one task, satisfaction, or completion. The result is a cross-surface topology of the emerging search journey and a discipline for distinguishing observed containment from inferred resolution.
Introduction. A person wants to understand a medical result, choose a car seat, or follow a breaking story. In the model that has organized information-retrieval research for decades, they begin in uncertainty, issue a query, scan a ranked list, reformulate, gather across sources, and synthesize, an iterative multi-step episode instead of a single lookup [4, 14, 18]. Conversational AI is now inserted into that episode. A common framing of what it does is the answer engine: a prompt goes in, a synthesized answer comes out, the episode ends. That framing has an appealing empirical signature. If a clickstream shows no onward search or pageview after an assistant response, it is natural to call the episode finished. But this reading depends on two decisions that do more work than they appear to. First, it treats prompts and responses as non-actions, so a provider session with several exchanges looks the same as one response. Second, it treats a lack of outward web activity as evidence about the information need rather than as evidence only about the observed surface. Counting conversational events is necessary; calling them one task or a successful resolution requires additional semantic evidence. We take a more structural view. The familiar observable shape of search begins with a query and moves outward: results, documents, reformulations, and synthesis. Conversational AI introduces another place where that work can occur. People can begin with the assistant, arrive there after encountering the web, remain inside dialogue, or move repeatedly between surfaces. The consequential question is therefore not simply whether an assistant replaces a query. It is where the assistant sits inside the broader observed journey. Two surfaces carry this AI. Search-embedded AI—Google’s AI Overviews and AI Mode—renders a synthesized answer inside the results page, with no session separable from the query it sits in. A standalone conversational assistant—ChatGPT, Claude, Perplexity, or Gemini’s own surface—is a destination the user navigates to, generating a session with its own surrounding web context. The before-and-after web context we measure is therefore defined only for the standalone case. We study the standalone surfaces; the in- SERP surface has its own measurable behavior, such as reduced onward clicking when a synthesized answer appears [19], that a within-SERP instrument rather than ours is suited to observe, and we treat it as a companion object requiring a different measurement design. We call a difference in that distribution recomposition: information-seeking activity is arranged differently across dialogue, search, and browsing. This is a descriptive claim, not a causal one. The comparison cannot separate what assistants change from the kinds of needs people choose to bring to them. Using an opt-in panel that links assistant activity to the same panelists’ observed search and browsing, we ask a question that classic web logs cannot answer and assistant-only corpora cannot either: where does external web activity fall around assistant use, and how does that differ from where it falls around conventional search? Our thesis is:
Search tends to open the observed journey toward content, while assistant use more often follows prior web activity. Measured with one construction on both anchors in the same users, the two shapes differ in direction and in containment, and no single “answer then exit” pattern describes the observed system. We make this concrete twice with one construction. First, a topology of AI-containing temporal sessions built from the position of assistant events relative to the web (Figure 2). We use observable labels—contained, AI-first, AI-last, and bridge/interleaved—because the trace tells us where events sit, not what they caused or whether they share one task. Second, the identical construction centered on conventional search in the same users’ non-assistant sessions (§6): the old shape, measured with the new instrument, so the two shapes can be compared like for like (Figure 1).
Contributions. (1) A method and unit: cross-surface session reconstruction that counts prompts and responses as first-class events alongside searches and pageviews, from a same-user panel, together with the direct positional marginals and the gap, coverage-density, and time-to-next-event diagnostics that expose which conclusions depend on session construction (§3, §4, Appendix A). (2) A full-session topology of AI-containing sessions, contained, AI-first, AI-last, and bridge/interleaved, reported both session-weighted and user-weighted with user-clustered intervals, stable in its AI-first and AI-last shares while the contained/bridge boundary moves with the segmentation gap; inside containment, a construct-validity decomposition shows that multiple captured responses do not by themselves establish one dialogue or successful resolution (§4, §5).
Related work. Information seeking is an episode, not a query. Classic models cast search as an affective and cognitive process from uncertainty toward focus instead of mechanical retrieval [14, 15], often beginning from an anomalous state of knowledge [4] or a sense-making gap [8, 27]. Exploratory search rejects the single-query model [18], strategies vary within one episode [3], the session rather than the query is the natural unit of analysis and evaluation [13], queries are reformulated as the user learns [9], search itself is a form of learning [22, 26], and foraging accounts describe how seekers trade cost against value in deciding where to look [20]. Conversational search formalizes the multi-turn answer-then-continue loop and treats follow-up turns as meaningful actions [21]. We treat the task episode as the conceptual unit but the inactivity-defined temporal session as the observed proxy; this distinction is central to the paper’s claim discipline. Inactivity-defined sessions are a longstanding measurement compromise whose thresholds are heuristic Taxonomies of search, by intent and by surface. Web search has long been classified by intent—Broder’s navigational, informational, and transactional split [5], recently revised for the generative era into knowledge-, guidance-, and output-seeking intents that deliberately span both search engines and AI chatbots [16]. Conversational information seeking is itself defined by multi-turn dialogue rather than by where a surface sits [29], and the “answer engine” framing names the synthesized-answer output without distinguishing its locus [24]. Our cut is complementary and measurement-driven: a search-embedded surface, whose answer has no session separable from the query it sits in, versus a standalone surface that generates a session with surrounding web context. The before/after topology we measure is defined only for the latter, which is why our instrument studies the standalone surface and leaves the embedded one to a within-SERP design.
Does AI displace search, and where does behavior go? Information seeking is among the most common conversational-AI uses [7], and a large same-user before/after study finds no drop in search usage after assistant adoption [23]. Aggregate trend claims are hard to identify because adoption coincides with activity bursts [2] that inflate volume comparisons. A methods companion on this panel makes the obstacle concrete: assistant use is itself timed by the user, so known-null timestamps drawn at comparably active moments reproduce much of the apparent post-event search lift [11]. We therefore avoid a volume verdict and study episode composition instead.
Method. Panel and measurement instrument. An opt-in cross-surface research panel, whose members consent to have their device activity metered for research, covers February 2026 in the United States and Great Britain. The provider’s metering software records two browser-level event streams for the same user. The pageview stream records user-facing page visits with a timestamp, the host-level The topology. For each AI-containing temporal session we locate the assistant span (first to last prompt/response event) and ask where external web steps fall relative to it: before the span, after it, or between assistant events. This yields four mutually exclusive temporal classes (Table 2): assistant-contained (no external step anywhere), AI-first (external only after the span), AI-last (external only before the span), and bridge/interleaved (external on both sides, between assistant events, or coincident with an assistant event). Between-only sessions belong to bridge but have nothing before the first assistant event, so before/after marginals are estimated directly rather than reconstructed by adding classes.
Four discretionary choices. Four choices are genuinely discretionary, and we fix defensible defaults instead of hiding them. (i) The 30-minute inactivity gap defines temporal-session boundaries; it is the conventional web-sessionization default and the middle of the sensitivity range we report at 15, 30, and 60 minutes (Table 3). (ii) Provider session_id defines an assistant session; it is a recorded grouping key, and the cross-surface inactivity rule may combine several provider sessions or split one. (iii) Search is defined by full host against the canonical search-host list in the footnote above; the registrabledomain alternative and its measured effect are reported there. (iv) The primary population requires one pageview-active day, the least restrictive threshold of the ladder in Table 4; we report thresholds through all 28 days. Workbench-versus-seeking, topical continuity, satisfaction, and per-task splits require a separately governed content-validation pass and are not inferred here.
The search label is a construct choice. Because both the AI-side search marginals and the entire comparator rest on what counts as “search,” we recompute the load-bearing quantities under four defensible rules using only privacy-safe retained fields (Table 1). Requiring a captured key phrase is conservative (phrase capture can fail on genuine searches) and dropping the Yahoo portal roots isolates portal-homepage traffic. Every variant stays within a narrow band of the primary rule, none approaches the rejected registrabledomain rule’s inflated values (46.6% / 50.6% search-anywhere), and the comparator’s containment level and the within-user contrast move by well under the contrast itself.
Discussion. What differs. Search-centered sessions are contained in 21.1% [20.8, 21.4] of cases session-weighted and 19.5% [19.2, 19.8] user-weighted, against 37.8% [36.9, 38.7] and 34.1% [33.5, 34.7] for assistant-centered sessions. The one-sided classes are search-first 15.4% [15.3, 15.5] / 17.9% [17.6, 18.1] and search-last 9.8% [9.7, 9.9] / 9.2% [9.0, 9.3]; bridge/interleaved is 53.7% [53.3, 54.0] / 53.5% [53.1, 53.8]. Content appears somewhere in 78.9% [78.6, 79.2] / 80.5% [80.2, 80.8] of search-centered sessions, before the first search in 44.2% [43.9, 44.5] / 46.7% [46.4, 47.1] and after the last in 54.2% [53.9, 54.5] / 60.2% [59.9, 60.6]. The contained/bridge boundary moves with the inactivity gap here too (contained 25.0% at 15 minutes to 16.4% at 60; interval half-widths at most 0.4 points, intervals in the The mirror. The clearest structural difference is directional, and we state it like-for-like: content pageviews on both sides. Around search, content mass sits after the anchor: content follows the last search in 60.2% [59.9, 60.6] of sessions user-weighted against 46.7% [46.4, 47.1] before the first. Around assistants the asymmetry flips: content precedes the first assistant event in 45.5% [44.9, 46.1] of sessions against 38.9% [38.4, 39.5] after the last. The same reversal holds session-weighted (54.2% [53.9, 54.5] versus 44.2% [43.9, 44.5] around search; 42.6% [41.8, 43.3] versus 36.5% [35.8, 37.3] around assistants), holds when search events are added back to the assistant side (51.9% [51.3, 52.5] versus 42.7% [42.2, 43.3]), and holds within panelists contributing both session types: the peruser before-minus-after content gap is +7.0 [6.4, 7.7] percentage points around assistants and -13.5 [-13.9, -13.2] around search. The single cleanest recomposition estimand is the paired difference of these directions—(assistant before minus after) minus (search before minus after), bootstrapped as one quantity per user—which is +20.6 [19.9, 21.3] percentage points. Descriptively, the search span more often precedes its content, while assistant spans more often follow observed web activity. Search-centered sessions also interleave more (bridge 53.5% [53.1, 53.8] versus 37.1% [36.5, 37.7]): the query-click-query loop is a more tightly alternating structure than dialogue, which concentrates consecutive assistant events. This is the observed positional grammar of the two anchors.
Within the same user. Because assistant adopters may simply browse differently, the cleanest contrast conditions on the person. Among panelists contributing both session types in the month, the within-user difference in contained shares (assistant-centered minus search-centered, equal-user weighting) is +13.0 [12.5, 13.6] percentage points. This conditions on the person, not on what the person brought to each surface; the task mix is what the next construction takes up.
Composition, not just position. The shapes also differ in what fills them. Composition is summarized session-weighted: within-session event shares and switch counts are averaged over pooled sessions with user-clustered intervals (surfaces are assistant, search, and content; a switch is an adjacent surface change in the time-ordered stream; durations are pooled medians). Within assistant-centered sessions, assistant events make up 59.7% of steps, search events 8.8%, and content pageviews 31.5%; the median session lasts 22 minutes and switches surface 5.2 [5.0, 5.4] times on average. Searchcentered sessions are 48.6% search and 51.4% content by steps, with a median duration of 9 minutes and 4.0 [3.9, 4.1] surface switches. These are event-share summaries of observed activity, not time-use estimates; dwell is observed only for pageviews.
Conclusion. The familiar query-to-content sequence now coexists, in the same panelists, with temporal sessions in which people bring prior browsing into dialogue, remain inside dialogue, begin from dialogue, or move repeatedly between the assistant and the web. Across two months, the old and new anchors show a consistent directional contrast: search tends to open the observed journey, while conversational assistants sit deeper inside it. That is the new shape of search. It is not a shorter version of the old one, and the data do not establish that assistants caused it. It is a measurable recomposition of where observed informationseeking activity occurs—and a reason to study dialogue, search, and browsing as parts of one cross-surface system.
Limitations. First, scope of claim: we characterize temporal-session composition and make no causal-volume claim. The topology does not identify demand created or destroyed, which an observational panel cannot recover against burst-selection [2, 11]. Second, task identity: temporal proximity does not establish semantic continuity. Concurrent needs can merge, one need can split, and provider sessions are not validated task labels. Two construction details cut the same way and are worth naming. Assistant hosts outside the canonical list (for example Grok, DeepSeek, Meta AI, you.com) are not recognised as assistant surfaces, so their pageviews stay in the web stream and count as ordinary web activity; a session that moves from a recognised assistant to an unrecognised one is therefore read as AI-last or bridge rather than contained. And the extract admits assistant events whose captured content is between 2 and 30,000 bytes, so an unusually long response is dropped from the stream, which can move the last observed assistant timestamp earlier and reclassify a following pageview from “between” to “after”. Both push the same way – against containment and against the directional contrast – so the reported figures are the conservative side of each.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does AI adoption reshape collaboration patterns in knowledge work?- Do early Muse users represent genuine demand for personal assistants or platform lock-in?
- What informal learning opportunities vanish when GenAI use stays hidden from colleagues?
- Which AI news assistant performs better than the others in this study?
- Do other AI assistants perform similarly on voter election questions?
- Does ChatGPT displace search engines or question-and-answer platforms?
- Does chatbot use for schoolwork reduce students' critical thinking skills?
- Do users click links within AI summaries or end sessions instead?
- Do searchers prefer clarity about AI involvement when viewing search overviews?
- How much do AI Overviews currently appear in Google search results?
- Do assistant sessions interleave with web content differently than search sessions do?
- Do AI-generated articles now dominate search results as organic traffic declines?
- How do publishers distinguish between search crawling and AI training requests?
- When do AI overviews beat conversational chat for answering user questions?
- Which search queries trigger AI summaries most often on Google?
- Are zero-click searches rising because of AI answer summaries?
- Does effort reduction during search affect how deeply people understand topics?