What Do Large Language Models Know? Tacit Knowledge as a Potential Causal-Explanatory Structure

Paper · arXiv 2504.12187 · Published April 16, 2025
Domain Specialization in LLMs

Abstract It is sometimes assumed that Large Language Models (LLMs) know language, or for example that they know that Paris is the capital of France. But what—if anything—do LLMs actually know? In this paper, I argue that LLMs can acquire tacit knowledge as defined by Martin Davies (1990). Whereas Davies himself denies that neural networks can acquire tacit knowledge, I demonstrate that certain architectural features of LLMs satisfy the constraints of semantic description, syntactic structure, and causal systematicity. Thus, tacit knowledge may serve as a conceptual framework for describing, explaining, and intervening on LLMs and their behavior.

Introduction. Although large language models (LLMs) are generally trained on next-word prediction tasks, these systems have exhibited impressive performance in the generation of seemingly human-like, coherent texts (e.g. ChatGPT (OpenAI 2022)). Yet, we do not know how these models work and what they learn from their data. Like other types of deep neural networks, LLMs are opaque and suffer from the black-box problem: during training, LLMs learn a highly complex function with distributed representations and we cannot simply “look inside” to determine how they work (Burrell 2016; Creel 2020). Because of this, it is difficult to evaluate their potential linguistic and cognitive capacities (Shanahan 2023).

The field of explainable AI (XAI) aims to solve the problem of opacity by developing explanations through mathematical techniques that show how or why a network makes a certain decision (Zednik 2021). One specific question that has been asked is whether LLMs learn any kind of rules or algorithms, akin to symbolic rules in classical systems (Olah et al. 2020; Pavlick 2023). Uncovering such rules would help explain how AI systems work, facilitate the evaluation of potential cognitive and linguistic capacities, and potentially allow for interventions in the model internals to update the system’s predictions.

Related work. In the particular case of LLMs, it has been suggested that they do not just perform next-word prediction based on surface statistics, but that they learn to follow symbolic rules (Pavlick 2023), develop linguistic competence (Mahowald et al. 2024), or even acquire a kind of knowledge (Meng et al. 2022; Yildirim and Paul 2024). Indeed, with respect to syntax, LLMs seem to successfully represent grammatical rules, and apply formal linguistic structure (Linzen and Baroni 2021; Mahowald et al. 2024). Less attention has been paid to semantics, but some very recent studies claim to have identified representations of world models (Li et al. 2023), implicit meaning (Hase et al. 2021), and facts (De Cao, Aziz, and Titov 2021; Meng et al. 2022). Although it is clear that many of these previous discussions are concerned with what LLMs “know”, it remains difficult to conceptualize exactly what this “knowledge” actually consists in. In this paper, I will focus on semantics, and in particular, LLMs’ knowledge of semantic rules considered as representations of facts or other propositions that play a causal role in the system’s Specifically, taking inspiration from debates between symbolic and connectionist AI in the 1980s and 90s (e.g. Clark 1991; Fodor and Pylyshyn 1988), I propose that tacit knowledge, as defined by Martin Davies (1990), provides a suitable way to conceptualize semantic knowledge in LLMs. Tacit knowledge, in this context, refers to implicitly represented rules that causally affect the system’s behavior. As connectionist systems are known not to have explicit knowledge, tacit knowledge provides a promising way to characterize and identify meaningful representations in the model internals. Moreover, if representations of knowledge can be successfully identified, these could be used both to explain how the model works and to potentially change the behavior of the system by overwriting or otherwise changing this knowledge.

​ This paper is not the first to propose tacit knowledge for connectionist neural networks. As mentioned previously, Davies (1990) originally developed the account of tacit knowledge as a way to attribute knowledge to connectionist systems in the absence of explicit rules. Nevertheless, Davies argues that connectionist neural networks do not actually meet the constraints needed for attribution of tacit knowledge. As such, my contribution goes beyond Davies by applying tacit knowledge to contemporary transformer-based LLMs and showing that these can in fact meet the constraints. More recently, Lam (2022) has argued that the identification and description of tacit knowledge should be a target for explainable AI, but does not demonstrate that neural networks can actually have such tacit knowledge.

Method. 3. Davies’ Account of Tacit Knowledge 3.1 Knowledge in neural networks Following Davies, I will motivate the focus on tacit knowledge specifically by first introducing two alternative notions of knowledge of rules: explicit representation as a strong notion of knowledge, and mere conformity as a weak notion of knowledge.

Symbolic AI systems such as expert systems like Cyc (Lenat and Guha 1989), are often said to have explicit knowledge of rules, in the sense that they have a pre-programmed knowledge base containing representations of rules that are invoked during processing.

These rules might be explicitly represented in the system’s source code, in which case an external observer might be able to “read off” the rules given sufficient background knowledge, access, and technical expertise. Alternatively, a system might make use of a database containing rules or logical statements. While this database might not be accessible to an external observer, any rules or statements in the database can be considered explicit knowledge to the system.

Explicit knowledge plays a causal role in mediating transitions from inputs to outputs. When processing a particular input, the system needs to identify and retrieve the appropriate piece of knowledge, which then plays a causal role in the transition to a particular output. For example, in a question-answering system, this knowledge could be a statement of the form “Paris is the capital of France”. When the model is queried with an input “What is the capital of France?”, it returns “Paris” because of this explicitly represented knowledge. Thus, given that a system is known to have explicit knowledge, it is possible to explain its behavior in virtue of the system having this knowledge.

​ Connectionist systems like neural networks, on the other hand, cannot be said to have explicit knowledge of rules (see also Fodor and Pylyshyn 1988). In contrast to symbolic systems, neural networks do not invoke rules represented in a knowledge base or database while processing inputs. Instead, neural networks are optimized using machine learning techniques to perform a particular function. This means that the way an input is processed is determined by the weights of the network, rather than by explicit rules stored elsewhere in the network. As the weights of individual nodes generally do not correspond one-to-one to rules due to distributed representation (Hinton, McClelland, and Rumelhart 1986; Van Gelder 1992), neural networks cannot be said to have knowledge of rules in the strong sense.

Instead, neural networks can be said to conform to rules, or have weak knowledge of rules (Davies 1990). Knowledge of rules in this sense does not require the system to represent rules in any way, but only requires that the behavior of the system—in terms of input-output transitions—matches what is described by the rule in question. Imagine we have a network that can correctly classify images of cats and dogs. Given that it successfully performs this task, the network could be said to conform to a rule “pointy ears → cat” and “floppy ears → dog”, but it could also and alternatively be said to conform to the rule “slitted pupils → cat” and “round pupils → dog”. This kind of indeterminacy shows the limitation of only attributing knowledge of rules in the weak sense: one network might conform to multiple sets of rules (Kripke 1982), or multiple networks might conform to the same sets of rules while relying on wildly different underlying processing mechanisms (see e.g. Block 1981). To distinguish between these kinds of systems, an intermediate notion of knowledge is needed that characterizes knowledge of rules without requiring explicit representation.

3.2 Tacit knowledge as an intermediate notion of knowledge Davies’ account of tacit knowledge is meant to provide an intermediate account of knowledge, one that is more than mere conformity but less than explicit knowledge. Tacit knowledge, according to this account, refers to rules that are not represented explicitly, but that nevertheless describe causally-relevant structures that guide behavior (Davies 1990).

Discussion. 5. Do LLMs have Tacit Knowledge?

In the previous section, I have shown that transformer-based LLMs meet a weaker version of the syntactic structure constraint, which still allows for the attribution of tacit knowledge. The next step is to verify whether LLMs meet the constraint of causal systematicity. In this section, I analyze a recent example from the technical literature (Meng et al. 2022) and show that the representations of factual associations identified in this work could be considered causal common factors in Davies’ sense, providing compelling—albeit preliminary—evidence that at least some LLMs meet the constraints for tacit knowledge.

5.1 An example In their paper, Meng and colleagues investigate factual associations in LLMs, specifically where and how facts are stored in the model internals. Take the example of an LLM that correctly predicts “Paris” in response to an input like “The Eiffel Tower is located in”, “Berlin” to “The Brandenburg Gate is located in the city of”, etc. While it is often suggested that such a network “knows” the location of particular tourist attractions purely based on its behavior, Meng and colleagues investigate where the network stores such facts and how potential representations are used in the processing of the network.

Meng and colleagues approach this problem in two steps: first, localizing potential representations of facts, and then, verifying that these play a causal role in the observed behavior. To localize representations of facts in the network, Meng and colleagues use a method called causal tracing, which determines which activations, for example at which layer, are most important for a particular prediction. To test this, the activations of the network are corrupted such that the network no longer predicts the correct outcome. Then, activations are iteratively restored to the original value, until the network returns the correct prediction. Using this method, activations in the MLP layers are found to play an important role in representing factual associations.

Based on these results, Meng et al. suggest that facts are stored in certain MLP modules in the middle layers of the network. Moreover, in line with Geva and colleagues (2021), they suggest that MLP modules might function as so-called key-value pairs, which take inputs that represent a particular input and return outputs that reflect memorized properties of that input. For example, given an input “the Space Needle is in downtown”, “Space Needle” might be viewed as a key, which prompts the MLP module to recall a particular value which represents properties of the Space Needle, such as its location. This information is accumulated over multiple layers and finally leads to a particular output at the last layer.

In the second part of their study, Meng and colleagues verify whether MLP modules function as key-value pairs and, if so, whether they can be updated to change the model’s behavior. To do this, they introduce a method called rank-one model editing (ROME) which identifies and applies edits to an LLM’s weights, based on two assumptions. First, if the MLP module functions as a key-value pair, this key-value pair can presumably be updated to include new information. Second, if the key-value pair plays a causal role in the model’s behavior, updating the stored information should be reflected in the model’s output. As such, the goal is to calculate a new key-value pair that encodes a particular updated factual association, e.g. “The Eiffel Tower is in Rome”. The existing key-value pair is then replaced with this new key-value pair (see figure 4), after which the output of the network is evaluated for various prompts. The edit is successful if the output of the network changes to the newly inserted fact.

Conclusion. The contributions of this paper are twofold. As a methodological contribution, I argued that we can take inspiration from Davies’ account of tacit knowledge to conceptualize semantic knowledge in LLMs. More precisely, Davies’ account provides clear criteria that should be met in order to attribute tacit knowledge to a particular system. While Davies himself argued that connectionist systems cannot meet the constraint of syntactic structure, I argued that this constraint can be appropriately weakened to acknowledge the role of the embedding layer in current LLMs. Armed with this weakened constraint, LLMs can in fact be attributed tacit knowledge, provided that they meet the other constraints. As an empirical contribution, I evaluated the recent work by Meng and colleagues (2022) to argue that there is compelling preliminary evidence to suggest that some LLMs actually acquire tacit knowledge in this sense. While this evidence is as of yet preliminary, tacit knowledge could thus be a promising tool to guide further research into the internal causal processing of LLMs.

Notably, this paper fits into a broader literature suggesting that LLMs represent something akin to knowledge in the model internals (Hase et al. 2021; Li et al. 2023; Pavlick 2023). Provided this is supported by future work, this would have promising implications for explainable AI. In particular, the identification of such knowledge structures would not only provide a novel way to explain how these systems work, by identifying internal common causal factors, but also to improve the performance of these systems, by applying interventions to change the internal knowledge representation so as to e.g. counteract misinformation, hallucination, and bias.

In this context, a promising direction of future research is the notion of interventions. Interventions are often used to investigate the causal processing within neural networks and to update the behavior of the network for particular inputs. In both cases, interventions are given a causal interpretation.

Limitations. Despite the promising results, there are some reasons for caution. First, although the reported specificity and generalization are high, they are not perfect. As such, updates might affect unrelated input-output pairs or fail to generalize. Possible explanations are that some updates do not target a causal common factor, or that causal common factors are not perfectly delineated in neural networks with distributed representation. The latter is of particular concern, as networks with distributed representation often exhibit polysemanticity or superposition, meaning that the same nodes in a network are involved in different decision processes (Van Gelder 1992; Elhage et al. 2022). Interventions might then affect multiple predictions, or even lead to catastrophic forgetting of previously learned associations. Further empirical research should determine to what extent this is a problem for interventions in practice.

Second, the work by Meng and colleagues (2022) relies heavily on interventions to locate and identify representations of tacit knowledge. Such interventions have recently been criticized, however. Specifically, while replicating the results reported by Meng and colleagues, Hase and colleagues (2023) found that interventions in different locations have a similar efficacy to the ones applied in the original study.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can artificial systems establish authority in domains requiring expert judgment? Can AI systems achieve real improvement without external human feedback? Does AI assistance help or harm professional skill development? Does AI-assisted research sacrifice exploration breadth for productivity gains? Is embodied interaction necessary for language meaning and agency? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Why do language models hallucinate and how can we prevent it? Can AI systems participate in genuine communication or only simulate it? How do neural networks learn compositional structure from training? Do language models reason through disagreement or only accommodate it? Can mechanistic interpretability methods reliably reveal what models actually know? How reliably can language models perform causal versus temporal reasoning? Can readers reliably distinguish AI-written text from human writing? Can language models reason beyond surface pattern matching? Can LLMs distinguish between linguistic form and semantic meaning? How do philosophical assumptions about AI consciousness affect practical harms and design? What prevents LLMs from applying their reasoning knowledge to improve outputs?