Can LLMs Help Improve Analogical Reasoning For Strategic Decisions? Experimental Evidence from Humans and GPT-4
This study investigates whether large language models (LLMs), specifically GPT-4, can match human capabilities in analogical reasoning for strategic decision-making. Using a novel experimental design that requires source–target matching, we find that GPT-4 achieves high recall—retrieving all plausible analogies—but suffers from low precision, frequently applying incorrect analogies based on superficial surface features. Humans, by contrast, exhibit high precision but low recall, selecting fewer analogies yet with stronger causal alignment. These findings advance theory by highlighting matching—the evaluative phase of analogical reasoning—as a distinct and critical step that requires accurate causal mapping over and above mere retrieval of an analogue. While current LLMs excel at retrieving analogies, human cognitive superiority persists in mapping causal structures across domains. Error patterns further reveal that AI failures stem from superficial similarity detection, whereas human errors reflect subtler misinterpretations of causal logic. In aggregate, the findings suggest a potentially productive division of labor in organizations: AI can serve as an analogy generator, while human decision makers act as critical evaluators in the application of the most contextually relevant causal schemas to organizational problem solving.
Introduction. Analogical reasoning—where managers draw on solutions from past experiences when dealing with novel challenges—is one of the central pillars of strategic thinking (Gavetti, Levinthal, & Rivkin, 2005; Miller & Lin, 2015). Analogies serve as cognitive tools that enable the interpretation of complex or ambiguous situations by linking them to familiar precedents (Gentner, Holyoak, & Kokinov, 2001). They are helpful in a wide range of strategic choice situations, including decisions about market entry, acquisitions, business model innovation, and organizational turnaround (Gavetti, Levinthal, & Rivkin, 2005). Beyond guiding choice, analogies also structure how managers mentally represent their strategic environment. By highlighting certain features and suppressing others, analogies influence how opportunities and threats are construed, thereby shaping the contours of strategic problem framing (Gary, Wood, & Pillinger, 2012).
When Charlie Munger famously remarked, “You’ve got to have models in your head... and you’ve got to array your experience—both vicarious and direct—on this latticework of models”,1 he was arguing that analogies from another domain can act as a form of guidance (a model) in ambiguous and data-poor domains. However, his statement about the virtues of possessing a repository of analogies also highlights an important challenge: one must select the appropriate analogy from a broader set of possible candidates—a challenge known as the ‘matching problem’ (Cummins, 1992; Gentner, Rattermann, & Forbus 1993). This is especially critical in strategy, where decisions are high-stakes and often irreversible (Leiblein, Reuer, & Zenger, 2018). A poorly matched analogy can lead to flawed inference, misplaced confidence, and costly strategic errors.
Matching correctly—often in a one-shot setting with no opportunity for feedback from implementation—requires not just an ability to recall past experiences, but also sophisticated judgment in selecting the appropriate one. Without accurate matching, a large store of possible analogies may, in fact, be a disadvantage. While stories of successful analogical transfer abound in the business world, it remains unclear whether these are indicative of the general power of analogies or simply the visible survivors, overlooking instances where analogical reasoning may have misfired. For instance, while Kodak executives successfully modeled their film and camera business on the analogy to how Gilette sold razors and razor blades, they also considered digital images to be analogous to traditional film, a mistake they did not recover from (Tripsas & Gavetti, 2000). Wayne Huizenga successfully leveraged his experience in the funeral homes business, to consolidate local video rental businesses into Blockbuster. But eventually Blockbuster paid the price for mistakenly viewing the Netflix business model as an analog to its own DVD distribution model. Overall, how human decision makers identify good matches and filter out bad analogies is an important and yet poorly understood phenomenon (Blanchette & Dunbar, 2001; Gentner & Smith, 2013) Rapid developments in Large Language Model (LLM) technologies are of relevance to the challenge of effective analogical reasoning. At its core, the transformer architecture (Vaswani et al., 2017) of the deep learning models that underlie LLMs allows them to detect similarity and relevance between ideas expressed as text, i.e., between a given prompt and patterns embedded in their training data. However, similarity detection is necessary, but not sufficient, for effective analogical reasoning (Olguin, Tavernini, Trench, Ricardo, & Minervino, 2022). Even if LLMs can outperform humans in the retrieval of potential analogies based on similarity, their ability to map structural similarities effectively to produce good matches remains unproven. On the dimension of structural abstraction, particularly in the domain of verbally formulated reasoning, as is common in business contexts, there is scant evidence on the relative performance of LLMs and humans (Yuan et al., 2023; for non-verbal reasoning see Lewis & Mitchell, 2025; Camposampiero et al., 2025). In other words, are LLMs capable of performing analogy retrieval and matching, and how does their performance compare with that of humans? Can human decision makers benefit from analogical reasoning executed by LLMs? What are the complementarities between analogical reasoning by human and AI agents?
This paper directly investigates these questions by comparing the verbal analogical reasoning abilities of GPT-4 and humans in a business context. In a novel experimental setup designed to replicate and meaningfully extend—classic work on analogical transfer, we introduced a matching problem by pairing two source analogs with two target problems that reflect typical business contexts.
Related work. Analogical reasoning plays a vital role in strategic and organizational decision-making, particularly under conditions of novelty, ambiguity, and incomplete information (Miller & Lin, 2014). Managers frequently draw analogies from past experiences to frame new challenges (Gavetti, Levinthal, & Rivkin, 2005; Tripsas & Gavetti, 2000). Analogies also shape how managers perceive opportunities and threats (Gary, Wood, & Pillinger, 2012). In entrepreneurial settings, analogies serve to frame new ventures, drawing on familiar categories that guide investor expectations and strategic alignment (Navis & Glynn, 2011; Santos & Eisenhardt, 2009). Similarly, analogical reasoning facilitates capability reconfiguration in dynamic markets by transferring lessons from prior contexts (Helfat & Peteraf, 2003; Zollo & Winter, 2002). Strategy narratives often invoke analogies from war, sports, or ecosystems to make complexity more tractable (Cornelissen, Holt, & Zundel, 2011).
Research in cognitive psychology recognizes analogical reasoning to be a fundamental cognitive process that enables the transfer of knowledge across domains (Gentner, 1983; Holyoak & Thagard, 1989). Researchers have converged on two sub-processes to describe analogical reasoning: retrieval, where possible source analogies come spontaneously to mind, and mapping, where the candidate analogies are compared to the target problem at hand. Underlying both subprocesses are the concepts of mental representations and similarity between them. One can represent two problem domains using networks which capture the key concepts and the causal interconnections between them. The two networks may resemble each other in terms of the nodes used (superficial or surface similarity) or the pattern of relationships between the nodes (structural or deep similarity). Successful analogical reasoning depends on finding structural similarities between situations, such that knowledge of a solution in one situation (source) can act as a useful hypothesis about the second (target) (Gentner & Smith, 2013; Goldwater & Gentner, 2015).
For human decision makers, spontaneous analogical retrieval is often cue-dependent and superficial (Gick & Holyoak, 1980). Similarity detection can occur purely based on superficial features across domains. Each feature in the source domain “activate[s] memory representations of other situations that share that feature.” (Holyoak & Koh, 1987; p. 333). Structural mapping involves checking for structural similarity across source and target domains (Gentner, 1983) via the detection of correspondence between “causal relations in the two situations.” (Holyoak & Koh, 1987; p. 334).
Method. 3.1. The experiment:
The influential studies by Gick and Holyoak (1980; 1983), demonstrated that when participants were given a single source story (e.g., a general attacking a fortress) and asked to solve a target problem (e.g., a doctor treating a tumor), many could apply the source —especially when given a hint. But this one-to-one setup does not involve matching, as it does not test whether people can discriminate between multiple, competing potential analogies—some superficially similar, others structurally appropriate. In other words, the Gick and Holyoak studies illuminated how analogical use can be elicited (i.e., they are tests of retrieval), but not whether analogical matching can be competently performed under realistic cognitive conditions. In our studies, we incorporated a matching problem, where participants were shown two source stories and faced two target problems. Our experiments thus departed from the classical ‘radiation problem’ studies on two important aspects: i) Inducing matching complexity: In addition to the ‘radiation problem’ story (Story 1), we introduce an additional story in the source domain (Story 2). Mirroring this in the target domain, we expose the subjects to two problems (Problem 1 and Problem 2) instead of one. To solve Problem 1, Story 1 should serve as the appropriate analogy while Story 2 would act as a placebo.
Similarly, to solve Problem 2, Story 2 should serve as the appropriate analogy while Story 1 would act as a placebo. Inducing this complexity to the traditional studies represents the “matching problem” that is at the heart of effective analogical transfer—a process that involves both the retrieval and accurate causal mapping of the right analogue. ii) Humans vs. AI: We run the above set-up for human subjects and independent trials on an established LLM-based AI developed by OpenAI, GPT-4. The human subjects comprise Masters’ level students from a leading business school.
We retain S1—the radiation problem—from the original Gick & Holyoak studies (1983) to benchmark our results against prior findings. The corresponding target problem, T1, is a novel but structurally analogous scenario in the domain of operations and supply chain management. We design S2 and T2 as a new pair to capture a different type of reasoning challenge—one rooted not in domain-specific knowledge, but in recognizing survivorship bias as an abstract causal schema.
This bias is broadly applicable across managerial contexts, especially when making inferences from incomplete data (Mundt, Alfarano, & Milakovic, 2022). Based on the above logic, we present two stories as follows:
Story 1 (the Radiation Problem: Split and converge schema):
“Dr. Clarke, a seasoned radiation oncologist, observed with concern as his patient, Sarah, battled severe side effects from her ongoing radiation therapy. Sarah was just one of many patients experiencing the unintended consequences of treating cancerous tumors with high-energy radiation. Although radiation therapy had been a critical component in the fight against cancer for decades, it was not without its flaws. The primary challenge lay in delivering a potent dose of radiation directly to the tumor while sparing the surrounding healthy tissue. Too often, this delicate balance proved difficult to achieve, resulting in collateral damage to nearby organs and exacerbating the patient's suffering.
Dr. Clarke knew that a more precise and targeted approach was needed, one that could improve treatment outcomes and enhance the quality of life for patients like Sarah. Dr. Clarke came up with a new technique that involved delivering multiple beams of radiation at varying intensities, aimed at the tumor from different angles. This allowed the cumulative dose of radiation at the tumor site to be sufficient to destroy it, while the surrounding healthy tissue would receive a lower dose, thus minimizing damage.”
Story 2 (the Dolphins Story: An example of survivorship bias):
Discussion. Without a hint, AI fails to recognize or apply analogies, resulting in zero precision, recall, and F1 score—essentially complete absence of the evidence for analogical reasoning. Human subjects, on the other hand, demonstrate high precision but very low recall, indicating that they only claim to use analogies when highly confident, often missing to spot an opportunity for analogical transfer (under-diagnosis). When given a hint, AI performance improves dramatically in recall (up to 1.0) but suffers from over-application of analogies, leading to moderate precision and many false positives (over-diagnosis). Human subjects also improve with hints, particularly in recall, but still lag behind AI in overall analogy detection. Their under-diagnosis pattern persists, which suggests that hints are not sufficient to overcome their conservative approach. Thus, our results indicate that providing a hint helps both groups, but AI shows a more dramatic reaction, which introduces a new trade-off between correct and incorrect analogy application.
The findings suggest that AI and human decision makers could be complementary in analogical reasoning to solve strategic problems. Current AI systems, when given a hint, demonstrate high recall—they rarely miss opportunities to apply analogies—but are prone to overapplication; often suggesting analogies that do not apply. In contrast, human decision-makers show high precision, typically identifying analogies correctly, but suffer from conservative approach; frequently failing to recognize relevant analogies–even when prompted. These patterns suggest a potentially productive division of labor in organizational decision-making that requires analogical transfer: AI can serve as an analogy generator, surfacing potential analogical matches across diverse contexts, while human decision makers act as critical evaluators, exercising contextual judgment to assess which analogies are truly applicable.
4.4. Nature of misappropriation of analogy (Surface vs. Structural similarity) Figures 3 (a), 3 (b) illustrate the nature of incorrect analogical transfer—i.e., misappropriation of analogies—by GPT-4 and human subjects, respectively. On the unconditional set of responses (i.e., hint or without hint), we first code each response as an instance of surface and structural similarity between the source stories and target problems based on the definition of “surface” and “structural” match of Holyoak and Koh (1987), and Gentner and Smith (2013).
We define incorrect analogical transfer as instances where participants applied the radiation story to solve the HR problem or the dolphin story to solve the city factory problem. Within this misapplication, we categorize the responses based on whether the mistaken analogy transfer was initiated due to surface similarity (e.g., dolphin→sea (the common feature)→city factory with port) or a “wrong” causal structure with respect to our pre-defined “correct” DAGs. In some cases, AI drew analogies from both stories to solve a particular problem, using a mix of incorrect surface and structural analogies from either story (cases coded as “both”). Effectively, these cases may represent partially correct analogical transfer. However, the propensity for AI to use both stories also highlights the ‘demand effect’—i.e., the LLM’s propensity to find analogies from both stories just because those stories are available and must be utilized regardless of the degree or nature of the similarity (structural or surface/superficial).
Limitations. As with any experimental study, certain limitations constrain the scope of our findings and point toward promising directions for future inquiry. Our human participant pool consisted of business school students—a common proxy for boundedly rational decision-makers in strategic cognition research, but nonetheless a sample of strategy novices. While appropriate for isolating cognitive mechanisms, replication with senior executives or cross-cultural cohorts would enhance the external validity of our claims. The task environment itself was deliberately simplified: participants were asked to select between two analogical sources for each of two target problems.
While this structure enabled tight control over the matching challenge, it does not fully reflect the complexity of real-world strategic decision-making, where the analogical search space may
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Does scaling reasoning capability create fundamental tradeoffs in control and reliability?- What explains the demand effect when available analogies get over-applied?
- How does task simplification affect analogical reasoning patterns?
- How sensitive is analogical reasoning emergence to training data and scale?
- Can transformers abstract relational structure without explicit symbolic machinery?
- How do transformers compose multi-step reasoning across different domains?
- Do causal rules enforce robustness that statistical patterns alone cannot maintain?
- What architectural features enable counterfactual reasoning in world models?
- Why do causal graphs alone fail to capture human reasoning processes?
- What makes causal belief networks more auditable than prompted personas?
- How do humans use associative reasoning without causal connections?
- Can causal models be extended to include non-causal cognition?
- What makes causal explanations stronger anxiety predictors than counterfactuals or dissonance?