Reimagining Legal Fact Verification with GenAI: Toward Effective Human-AI Collaboration

Paper · arXiv 2602.06305 · Published February 6, 2026
AI at Work

Fact verification is a critical yet underexplored component of nonlitigation legal practice. While existing research has examined automation in legal workflow and human-AI collaboration in highstakes domains, little is known about how GenAI can support fact verification, a task that demands prudent judgment and strict accountability. To address this, we conducted semi-structured interviews with 18 lawyers to understand their current verification practices, attitudes toward GenAI adoption, and expectations for future systems. We found that while lawyers use GenAI for low-risk tasks like drafting and language optimization, concerns over accuracy, confidentiality, and liability are currently limiting its adoption for fact verification. These concerns translate into core design requirements for AI systems that are trustworthy and accountable. Based on these, we contribute design insights for human-AI collaboration in legal fact verification, emphasizing the development of auditable systems that balance efficiency with professional judgment and uphold ethical and legal accountability in high-stakes practice.

Introduction. Fact verification is a critical process in non-litigation 1 legal practice, ensuring that all information supporting contractual commitments, regulatory filings and advisory decisions is not only accurate and complete, but also consistent and defensible [54, 77, 82, 83]. Highstakes business transactions [52] and compliance processes [9] depend on the careful verification of factual information. A single oversight, such as failing to detect an undisclosed beneficial owner or misinterpreting a cross-border regulatory requirement, can lead to severe consequences including financial penalties, reputational harm, and prolonged disputes [29]. However, verification is far from a straightforward checklist. It is a complex and interpretive process that requires lawyers to synthesize a large amount of information from diverse sources [71], such as client-provided materials, official registries, and third-party reports, while reconciling inconsistencies and evaluating the credibility of each source [57]. Moreover, as new documents and contextual details emerge during the course of legal services, lawyers revisit and refine their factual understanding. Consequently, rather than a static confirmation of individual facts, fact verification operates as a dynamic sensemaking activity that combines retrieval, comparison, evaluation, and synthesis to construct a coherent and defensible narrative of the matter [18]. Such tasks are cognitively demanding and operationally fragmented. Lawyers must reconcile voluminous, often contradictory information under significant time constraints [40]. Existing legal technology solutions [27, 64, 70] provide partial relief by automating specific subtasks [58]. For example, Westlaw [88] facilitates legal research across jurisdictions, and analytics tools such as Luminance [51] accelerate due diligence by detecting clause-level inconsistencies. However, these tools operate within structured, task-specific contexts and fall short in supporting crossdocument reasoning or interpretive synthesis. They can highlight a missing warranty clause in a contract, but they cannot determine that an ownership disclosure in one agreement conflicts with shareholder information contained in another document. Such limitations leave lawyers responsible for the interpretive work required to maintain factual coherence across dynamic and distributed information landscapes [12]. Currently, Generative AI (GenAI) is being explored for implementation in professional fields such as medicine [37, 100], finance [44, 78], and scientific research [4, 35] through diverse technical approaches, including fine-tuning large language models (LLMs) [59, 87], integrating domain-specific knowledge retrieval [50], and optimizing task instruction design [102]. Unlike traditional automation, GenAI with advanced capabilities in language understanding, contextual reasoning, and multi-turn interaction, can potentially assist fact verification [1, 28]. However, its application to fact verification, the most judgment-dependent and high-stakes aspect of non-litigation practice, remains largely unexplored. Integrating GenAI introduces critical Human-Computer Interaction (HCI) questions: How should GenAI engage in fact verification without undermining professional accountability? How can interaction designs support interpretive flexibility, sensemaking, and trust while mitigating over-reliance? To address this gap, we conducted semistructured interviews with 18 experienced non-litigation lawyers. Specifically, we focus on the following research questions: RQ1: How do lawyers conduct fact verification in nonlitigation practice, and where does GenAI fit in these workflows? RQ2: What cognitive and operational challenges influence lawyers’ willingness and ability to use GenAI for fact verification? RQ3: Given these work practices and expectations, what forms of human-AI interaction can best support reliable and accountable fact verification? By examining these questions, our study offers an integrated view of how lawyers navigate the ongoing construction of factual understanding and where current tools fall short in supporting this work. The interview findings illuminate the practical realities behind fact verification by showing how lawyers synthesize evolving information, reconcile inconsistencies, and maintain defensibility under time pressure. This reveals how GenAI might meaningfully contribute to or potentially complicate the lawyers’ efforts to improve their work efficiency. In doing so, this research advances a deeper understanding of what reliable and accountable fact verification requires in practice, and outlines how future HCI can better support the interpretive and judgment-dependent nature of non-litigation legal work.

Related work. Legal technology, or LegalTech for short, refers to tools and systems that leverage technologies such as artificial intelligence (AI), big data, and cloud computing to optimize legal business processes [5, 39, 94]. Its core objective is to improve the efficiency of legal work, reduce costs, and minimize human errors through technical means, covering the entire business chain, from contract management and case retrieval to compliance review [91]. Early systems primarily focused on single-task automation, such as clause comparison, semantic search for statutes, and document drafting. The focus on the automation of individual tasks has yielded tangible results in various legal domains [45]. In the field of contracts, previous works show that combining information design and computer codification can enhance communication, participation, and usefulness across the entire life-cycle of contracting [8]. Kozlova et al. [46] elaborated the advantages and disadvantages of using both smart contracts and classic contracts in contractual practice. Their research suggests that standardized contracts can be completely replaced by smart contracts. Within retrieval and analysis, intelligent systems enable semantic matching for case and regulation searches, recommend similar cases, and analyze judicial trends of such cases [55, 68]. For example, Nguyen et al. [60] developed Attentive Convolutional Neural Network (Attentive CNN) and Paraformer, using deep neural networks with attention mechanisms. These tools proved better performance in terms of retrieval performance across datasets and languages. For documents and processes, automation extends to filling legal documents, such as complaints and evidence lists, as well as handling repetitive tasks such as format review and information extraction [25, 81]. This single-point automation provides significant value by reducing the mechanical labor of lawyers on low-value tasks, such as reviewing standardized contract clauses [42]. Fact verification remains inherently dependent on nuanced professional judgment, as it requires contextual understanding, and legal sufficiency evaluation. However, most LegalTech tools are still confined to mechanical or highly structured subtasks, leaving complex, judgment-dependent activities such as fact verification to human professionals [66]. For example, while AI tools can detect missing data or flag inconsistent entries in contracts, they lack the capability to evaluate the factual accuracy of complex disclosures, making them insufficient for tasks requiring legal judgment and accountability [53, 69]. This gap motivates investigation into whether and how GenAI can assist verification without undermining professional standards.

Method. We recruited 18 non-litigation lawyers, each with at least one year of professional experience in non-litigation practice. Lawyers with less than one year of experience and those whose primary work involved litigation were excluded from the study. Recruitment was carried out through snowball sampling. The 3rd author shared a recruitment message through personal social media, and initial respondents invited colleagues from their professional networks to participate. The final sample included 7 male and 11 female lawyers who worked in practice areas such as corporate transactions, crossborder compliance, and contract review. We intentionally included participants without prior experience in GenAI in order to capture a wide range of verification practices, existing challenges, and expectations regarding the role of GenAI in legal fact verification. Each participant received a 200 RMB JD E-gift card as compensation for their time. Table 1 details the characteristics of the participants.

We conducted semi-structured interviews remotely using Tencent Meeting to provide flexibility for participants’ schedules. Prior to participation, they were informed about the purpose of the study, the voluntary nature of participation, and how their data would be handled. They were also assured that they could withdraw at any time without penalty. Each interview lasted approximately 45 to 60 minutes and was conducted by the same author to maintain consistency and ensure familiarity with the entire dataset. All sessions were audio-recorded with permission and automatically transcribed using Tencent Meeting’s transcription service. The research team manually reviewed and corrected transcription errors to ensure accuracy. The interview questions were designed to align with our research questions while allowing participants to elaborate on their experiences in their own terms. The first part focused on fact verification in practice, probing how participants defined verification, what steps they followed, and what cognitive or operational challenges they encountered. The second part examined lawyers’ perceptions of GenAI, including their awareness of such technologies, perceived opportunities and risks, and scenarios where GenAI could potentially support verification tasks. The final part invited participants to envision collaboration with GenAI, exploring preferred interaction modalities, levels of control, and expectations for trust and reliability. The full list of questions is provided in the supplementary materials.

Two authors applied thematic analysis following Braun and Clarke’s six-phase framework [11], with all identifying details anonymized due to the sensitivity of legal work. Both authors began with the familiarization phase, reading the transcripts multiple times to gain a holistic understanding of the data. During the initial coding phase, each author independently coded the transcripts in a shared Microsoft Excel sheet organized by participant ID, interview prompt, verbatim excerpt, descriptive code, and analytic memo. We applied descriptive codes to segments relevant to the research questions and added new codes whenever excerpts did not align with existing ones. Both coders progressed in participant order and occasionally revisited earlier transcripts when later excerpts suggested that previous coding required refinement. After roughly every 4 interviews, we compared our coding decisions, discussed ambiguous excerpts, clarified overlapping labels, and resolved disagreements by returning to transcript context and consulting memos, which enabled us to maintain consistent coding rules throughout the process. This process was inductive, allowing codes to form from the data rather than being predetermined. Once initial coding was complete, we began grouping related codes into broader analytic categories, such as “document generation with GenAI,” “human verification and ultimate accountability,” and “external constraints from client.”

Discussion. Efficiency in legal fact verification cannot be understood as simply accelerating routine tasks. The findings in Section 4.3.1 reveal that what lawyers seek is a form of automation that strengthens the evidentiary infrastructure of their work rather than replacing human reasoning. Because fact verification in non-litigation practice involves reconstructing a defensible account of “what is true” across fragmented registries, client disclosures, interviews, and onsite observations, efficiency gains depend on whether GenAI can reduce uncertainty at the earliest stages where facts are gathered, organized, and stabilized. In this process, lawyers carry full accountability for every factual claim they endorse. As a result, automation is acceptable only when it reduces ambiguity without altering the chain of evidentiary interpretation. Section 4.1.3 shows that the heaviest cognitive burden lies not in legal reasoning itself but in the upstream work of transforming unstructured information into verifiable elements. Tasks such as extracting shareholder changes from interview notes, comparing multiple versions of corporate records, or building chronological transaction tables require continuous attention to consistency and timestamp alignment. These preparatory steps determine which details can be safely carried into risk evaluation, disclosures, or due diligence deliverables. It reveals that while automation in many fields offloads repetitive tasks, in law, similar offloading risks inadvertently shifting epistemic responsibility. AI-generated summaries often appear convincing but provide little visibility into the selection or interpretation of information [86]. Once such content enters a working paper, verifying it requires retracing the system’s steps, often consuming more time than manual work [73]. Efficiency collapses not only due to inherent GenAI inaccuracy, but more fundamentally because its opacity obscures the source of error [41], thereby disrupting the evidentiary chain that lawyers must preserve, which in turn suggests automation’s true value lies not in removing simple work but in its potential to stabilize the compromised factual substrate. The inability to delegate extraction and organization due to accountability needs means that efficiency is achieved only when GenAI acts as an infrastructural intermediary, which directly reduces the cognitive load caused by fragmented data environments by producing structured intermediate artifacts with explicit provenance. This enables lawyers to begin their reasoning from a more reliable evidentiary baseline. Structured GenAI that can automatically align data across sources, flag timestamp discrepancies, or generate audit trails directly makes it possible for lawyers to shift their cognitive effort from reconstructing factual coherence to evaluating legal implications [56]. Meaningful automation depends on systems that first act as stabilizing forces for data fidelity, ensuring that the material lawyers interpret remains trustworthy. Only then can the boundaries of human legal responsibility be clearly defined. Ultimately, this reframes efficiency as an epistemic design problem, demanding that AI systems be engineered to support the evidentiary commitments of legal practice rather than merely its workflow timelines.

While GenAI offers opportunities to alleviate cognitive burden in fact verification, our findings in Section 4.1.2 show that professional expertise remains essential, particularly when lawyers must navigate uncertainty, reconcile fragmented evidence, and make defensible judgments.

Conclusion. In this study, we conducted a qualitative analysis of interviews with 18 legal professionals to explore their experiences, attitudes, and expectations regarding GenAI integration in legal fact verification. Our findings reveal that GenAI not only change the pace of fact verification but also influence how lawyers interpret information, form judgments, and manage responsibility within their workflows. Participants described both the benefits of reduced manual burden and the risks that arise when AI-generated content shapes early understanding of the fact or interrupts opportunities for developing professional intuition. These insights highlight the need for auditable AI systems that support transparency, preserve human oversight at critical decision points, and maintain conditions that allow legal expertise to grow. By grounding these implications in real practitioner experiences, our study offers guidance for designing AI tools that can be adopted responsibly in everyday legal work. More broadly, the work contributes to HCI research on expert–AI collaboration by showing how efficiency gains must be balanced with the practices that enable reliable professional reasoning.

Limitations. As with most interview-based research, our study has limitations that should be considered when interpreting the findings. Because the data relies on practitioners’ self-reported accounts, participants’ descriptions of their verification routines and GenAI use may diverge from actual practices, particularly in fast-paced non-litigation workflows. Although the interviews encouraged concrete examples and detailed reflection, discrepancies between articulated strategies and enacted behavior may persist. Future research could complement this work with observational studies to examine how G enAI-mediated verification unfolds in practice. In addition, while our discussion engages with prior work on GenAI in high-stakes domains, the analysis does not provide a comprehensive engagement with adjacent literatures such as health AI or other high-stakes professional domains. This limits the extent to which our findings are positioned within broader cross-domain debates in HCI.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does AI deployment reduce or exacerbate workplace inequality and income instability? How do hallucinated citations emerge in AI scholarly output? Does AI-assisted work increase total productivity or just shift time? Why does polished AI output gain credibility despite fundamental verifiability problems? How does AI adoption reshape collaboration patterns in knowledge work? What are the real-world consequences of AI citation hallucinations? How do educators verify student capability when AI can produce indistinguishable work? How do AI-exposed occupations change in employment, wages, and skills? How should human-AI contributions be measured, disclosed, and verified? Does AI assistance erode cognitive skills while inflating perceived competence? Can artificial systems establish authority in domains requiring expert judgment?