Seeing to Think? How Source Transparency Design Shapes Interactive Information Seeking and Evaluation in Conversational AI
Conversational AI systems increasingly function as primary interfaces for information seeking, yet how they present sources to support information evaluation remains under-explored. This paper investigates how source transparency design shapes interactive information seeking, trust, and critical engagement. We conducted a controlled between-subjects experiment (N=372) comparing four source presentation interfaces—Collapsible, Hover Card, Footer, and Aligned Sidebar—varying in visibility and accessibility.
Using fine-grained behavioral analysis and automated critical thinking assessment, we found that interface design fundamentally alters exploration strategies and evidence integration. While the Hover Card interface facilitated seamless, on-demand verification during the task, the Aligned Sidebar uniquely mitigated the negative effects of information overload: as citation density increased, Sidebar users demonstrated significantly higher critical thinking and synthesis scores compared to other conditions. Our results highlight a trade-off between designs that support workflow fluency and those that enforce reflective verification, offering practical implications for designing adaptive and responsible conversational AI that fosters critical engagement with AI generated content.
Introduction. Conversational AI systems are increasingly used as primary interfaces for information seeking, learning, and evidencebased writing. Unlike traditional search engines, these systems actively frame and summarize information through dialogue, shaping how users interpret evidence, assess sources, and decide when further verification is needed. This shift raises important concerns about how users form trust, how they recognize uncertainty and biases, and how they integrate multiple sources into their reasoning and task completion [43]. Prior research shows that interface transparency can strongly influence users’ mental models and reliance on algorithmic outputs [20, 32], while interaction design choices affect whether users critically reflect on system suggestions or accept them with minimal scrutiny [3, 8].
At the same time, studies of human–AI collaboration suggest that users often over-rely on fluent AI responses unless the Authors’ Contact Information: Jiangen He, The University of Tennessee, Knoxville, TN, USA, jiangen@utk.edu; Jiqun Liu, The University of Oklahoma, Norman, OK, USA, jiqunliu@ou.edu.
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org.
© 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. Manuscript submitted to ACM Manuscript submitted to ACM 1 arXiv:2601.14611v1 [cs.HC] 21 Jan 2026 2 Trovato et al. interface encourages deliberate engagement [10, 59]. Together, these findings highlight the importance of understanding interactive information seeking and evaluation as a core component of responsible conversational AI [59].
However, studying this process in conversational AI is particularly challenging because source use is mediated by interface design rather than direct document interaction. Source transparency is not only about whether citations exist, but also about how visible, accessible, and cognitively affordable they are during interaction. Research on algorithmic awareness shows that when system mechanisms are hidden, users may construct inaccurate explanations of how results are generated [12, 20]. Meanwhile, work on explainability and transparency demonstrates that explanations and documentation can fail to improve understanding, or even create false confidence or mistrust, depending on how they are presented [9, 45, 50]. In conversational AI systems, fluent language further blurs the boundary between generated content and external evidence, making it harder for users to distinguish, verify, and reconcile sources. As a result, citation interfaces may either support careful evaluation or subtly discourage it, depending on how they distribute attention and effort across reading, verifying, and writing activities [10, 70].
Motivated by these challenges, this work investigates how source transparency design shapes interactive information seeking and evaluation in conversational AI. We examine how different source display and access mechanisms (e.g. click, hover, listing) influence users’ exploration behavior, chat perception, and evidence use during writing tasks.
Our research questions focus on the interaction between interface design and information density, and on how this interaction affects critical engagement with information sources cited by system responses. The novelty of this work lies in treating source transparency as a behavioral and interactional phenomenon rather than a purely normative design principle [50, 68]. This study bridges research on human–AI interaction, algorithmic transparency, and critical reliance by grounding these discussions in fine-grained behavioral evidence [3, 9, 10]. In addition, the findings will offer concrete guidance and practical implications for designing citation and source presentation interfaces that support verification without overwhelming users and faciltiate evidence-based critical thinking and learning, contributing to more trustworthy and responsible conversational AI experiences [38, 44, 59, 70].
In this paper, we investigate the impact of citation presentation on user behavior and trust in AI-generated text. We guide our study with the following Research Questions (RQs):
RQ1: How do variations in source visibility and accessibility shape users’ information-seeking strategies and integration workflows?
RQ2: To what extent does source presentation design influence perceived transparency and trust?
Related work. 2.1 Conversational Information Seeking and Evaluation Conversational AI systems have transformed information seeking from a document-centered activity into a dialogic, natural-language-driven, and system-mediated process [74]. Instead of navigating ranked lists, users increasingly rely on AI agents to synthesize, prioritize, and narrate information through natural language interaction [34]. Prior work in human–AI interaction shows that such mediation shapes users’ judgments, confidence, and verification behaviors [3, 10].
From an interactive information retrieval perspective, information seeking is inherently iterative and situated [46], and conversational interfaces further blur the boundaries between retrieval, interpretation, and judgment [61, 75]. Generative responses compress multiple sources into a single narrative form, which can obscure uncertainty and disagreement while increasing perceived coherence and fluency [26, 70]. As a result, conversational information seeking introduces new challenges for how users recognize uncertainty, compare evidence, and decide when additional exploration is necessary, particularly in tasks that require critical evaluation rather than simple fact lookup [12, 22, 51, 68].
Evaluation in conversational information seeking therefore extends beyond factual correctness to include judgments of relevance, credibility, sufficiency, and argumentative support across sources [14]. Research on boundedly rational search and AI-assisted decision making shows that users rely heavily on heuristics and interface cues under cognitive constraints [4, 28], often defaulting to system suggestions unless interaction design introduces cognitive forcing or reflective friction [10, 11, 45, 63]. At the same time, studies of algorithmic awareness indicate that users frequently hold incomplete or inaccurate mental models of how conversational systems retrieve and generate information [20, 32, 35], which further complicates their ability to critically evaluate AI-mediated evidence [66]. Work on transparency and explanation mechanisms suggests that providing more information does not necessarily improve understanding, and can sometimes increase misplaced confidence or shallow verification [9, 50? ]. While prior studies have examined trust, transparency, and reliance in conversational and generative systems [12, 59], fewer have investigated how users behaviorally explore, revisit, and integrate sources during extended complex conversational tasks such as writing and argument construction. This gap motivates our focus on conversational information seeking and evaluation as an AI-faciliated interactional and interface-dependent process rather than a purely cognitive or perceptual outcome.
Method. 3.1 Interface Design To investigate how different source presentation formats influence user experience and chat experience, we implemented four distinct interface conditions (see Figure 1). Except the baseline interface, the four interfaces with source presentation presents the same source details (title, domain name with favicon, and text snippet), but each interface varies in how sources are displayed and accessed. Across all four conditions, participants can read sources by clicking on their badges or cards. They can also cite sources in their essay by dragging and dropping them directly into the essay editor to automatically insert a citation.
Sources (collapsible). Click to view the full list.
Hover over an inline citation to preview its source card.
The full source list appears at the end of each response.
Source cards appear in a rightside panel that stays aligned with the nearest inline citation and updates as scrolling.
Collapsible List Hover Card Footer Sources Aligned Sidebar Participant Task Sidebar (next to the Chatbot) Drag and Drop a inline Citation Badge to Cite Drag and Drop a Citation Card to Cite Type [4] to Cite Auto-generated reference list Fig. 1. Interface conditions 3.1.1 Collapsible Interface. In the collapsible interface condition, source badges appear inline within the text as clickable links. A collapsible “Used Sources” button displays above the message content. Participants can toggle this button to reveal the complete list of source cards.
3.1.2 Hover Card Interface. In the hover card interface condition, when users hover over a citation badge, a card appears displaying the source details. This design provides on-demand access to citation details, allowing users to quickly preview sources while maintaining their position in the text.
Manuscript submitted to ACM 6 Trovato et al.
3.1.3 Footer Interface. In the footer interface condition, all source cards are grouped at the end of each response message. All source cards are displayed by default without the need to expand, compared to the collapsible interface.
This approach mimics bibliography formatting, consolidating all reference information in a location.
3.1.4 Aligned Sidebar Interface. The aligned sidebar interface condition presents source cards in a persistent sidebar We use the four interface conditions to manipulate the source visibility and accessibility. Visibility refers to the explicit display of source details, while accessibility refers to the proximity of source details to the in-text citations. Based on these dimensions, the Collapsible List condition represents low visibility and low accessibility. The Footer condition provides high visibility but low accessibility. The Hover Card condition offers low visibility but high accessibility.
Finally, the Aligned Sidebar condition achieves both high visibility and high accessibility.
3.1.5 Citing Sources. Citations could be integrated into the essay via two methods: a drag-and-drop mechanism, allowing users to drag citation badges or cards directly into the editor, or a smart-typing feature, where typing a citation index (e.g., “[1]”) automatically converted the text into a linked citation. Once inserted, citations appeared in the editor as clickable, blue, bracketed links (e.g., [1]). Users could remove citations using standard text deletion methods (Backspace key), which treated the citation as a single unit. To assist with bibliography management, the system dynamically generated and updated a “Cited References” list below the editor, displaying the full details (title and URL) of all unique sources currently cited in the text.
3.2 AI Agent The AI agent was powered by Perplexity AI’s sonar-pro model, configured with a temperature of 0.7. The model utilized web search capabilities with medium search context size and user location set to ‘US’ (United States) to retrieve real-time information from the internet and generate responses with inline citations. To ensure high-quality, evidentiary support, the agent operated under a system prompt (see Appendix A) that defined its role as a research assistant.
Discussion. Scaling Critical Thinking with Information Density. A central finding of our study is that the is the discovery of that effect design can reverse the negative impact of information overload on information evaluation and critical thinking. In typical search and chat interfaces, increased information density often leads to cognitive fatigue and shallower processing [4, 10, 60]. It also refelects in our results: as the AI provided more citations, participants in the Collapsible condition (lower visibility and availability of sources) show decline in critical thinking scores when information density is high (Figure 4). The behavior patten tells the same story that designs with lower visibility and availability of sources can lead to lower engagement as conversation evolves (Figure 2) or less deep engagement (Table 4) . However, our Sidebar condition (higher visibility and availability of sources) reversed this trend: as the AI provided more citations, participants in the Sidebar group showed significantly improved scores in Synthesizing Multiple Sources and Source-Related Critical Thinking. We may attribute this to the role of the sidebar as an external working memory.
By presenting sources from the ephemeral, linear flow of the chat and placing them in a persistent, spatial layout, the Sidebar allows users to “offload” the cognitive burden of tracking evidence. Our findings echo prior work showing that transparency cues do not have uniform effects on user trust or understanding, but instead interact with task context and interface design [e.g. 15, 29, 32]. We also extend these insights by providing behavioral and critical-thinking evidence that transparency effects in conversational AI are strongly conditioned by information density and interaction layout rather than source presence alone.
The Trade-off Between Flow and Verification. Our behavioral sequence analysis reveals a distinct trade-off between interfaces that support “flow” and those enhance “verification”. The Hover interface was unique in supporting a seamless Cite →Write loop (Prob. = 0.393), allowing users to verify specific claims on-demand without breaking their drafting context. This resulted in higher scores for Evaluating Evidence Strength, as users could perform micro-verifications of individual facts. However, this efficiency came at the cost of synthesis. The Sidebar condition promoted a more disruptive Write →Explore loop, encouraging users to stop writing and scan multiple sources. While this friction likely contributed to the lower perceived usefulness ratings for the Sidebar if information density is low, it ultimately supported better holistic argumentation. This trade-off implies that there is no single optimal transparency design.
Instead, conversational AI systems may need interfaces tailored to different tasks and information densities, or adaptive interfaces—for example, using hover-based interactions during rapid drafting to maintain flow and transitioning to persistent sidebars during review or ideation to support synthesis and complex reasoning. This trade-off resonates with prior findings that interaction friction can function as a cognitive forcing mechanism that reduces over-reliance on AI while simultaneously increasing perceived workload [e.g. 9, 10]. This trade-off is further consistent with previous studies indicating that user trust calibration depends not only on explanation availability but also on how explanations are integrated into information interaction workflows [33, 40].
Conclusion. This work revealed a critical trade-off between interface designs that prioritize workflow fluency and those that support cognitive persistence. While low-friction mechanisms like hover cards facilitate immediate, on-demand verification during drafting, persistent layouts like the aligned sidebar act as external working memory, enabling users to sustain critical synthesis even as information density increases. Consequently, there is no “one-size-fits-all” solution for citation display. To foster responsible human-AI collaboration, future systems must move beyond static citations toward adaptive interfaces that balance the cognitive needs of reading, verification, and synthesis based on the nature of informationseeking tasks. Future research should further analyze the prompting process of participants and explore adaptive interfaces that dynamically shift between low-friction and persistent citation layouts based on real-time detection of user intent (e.g., browsing vs. critical analysis). Additionally, longitudinal studies are needed to determine if the benefits of “reflective friction” persist over time.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do writers navigate authorship and delegation with AI? Does AI assistance erode cognitive skills while inflating perceived competence? Why do confident AI outputs mislead human trust calibration?- Does mandatory AI disclosure in policy help or harm user trust over time?
- Does expressing emotion change how users trust an AI system?
- Can disclaimers alone prevent users from trusting AI outputs too heavily?
- Can transparent and aligned AI reduce consciousness attribution by users?
- Which interaction design changes most effectively prevent consciousness attribution?
- What responsibility do designers bear for consciousness attribution risk?
- How does understanding persistent journeys intensify both trust and privacy concerns?
- How does personalization increase trust while degrading clinical safety outcomes?
- Does transparency about AI use change how audiences trust the writing?
- How does the cultural reflex around advertising disclosure compare to AI disclosure?
- Can content-side interventions reduce AI persuasion where disclosure labels fall short?
- What threshold of skepticism does AI awareness actually create in audiences?