Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero

Paper · arXiv 2310.16410 · Published October 25, 2023
Correct but Not Understood

Abstract Artificial Intelligence (AI) systems have made remarkable progress, attaining super-human performance across various domains. This presents us with an opportunity to further human knowledge and improve human expert performance by leveraging the hidden knowledge encoded within these highly performant AI systems. Yet, this knowledge is often hard to extract, and may be hard to understand or learn from. Here, we show that this is possible by proposing a new method that allows us to extract new chess concepts in AlphaZero, an AI system that mastered the game of chess via self-play without human supervision. Our analysis indicates that AlphaZero may encode knowledge that extends beyond the existing human knowledge, but knowledge that is ultimately not beyond human grasp, and can be successfully learned from. In a human study, we show that these concepts are learnable by top human experts, as four top chess grandmasters show improvements in solving the presented concept prototype positions. This marks an important first milestone in advancing the frontier of human knowledge by leveraging AI; a development that could bear profound implications and help us shape how we interact with AI systems across many AI applications.

Introduction. Artificial Intelligence (AI) systems are typically treated as problem-solving machines; they can carry out the jobs humans are already capable of but more efficiently or with less effort, which brings clear benefits in several domains. In this paper, we pursue a different goal: treat AI systems as learning machines and demand from them to teach us the fundamental principles behind their decisions to extend upon and complement our knowledge. We can imagine many benefits of learning from machines. For example, while a system capable of producing a more accurate cancer diagnosis or effective personalised treatment than human experts is useful, transferring the rationale behind their decisions to human doctors could not only bring advances in medicine but also leverage human doctors’ strength and generalisation ability to enable new breakthroughs. There is a tremendous untapped opportunity across various domains where the capabilities of AI systems are reaching or exceeding those of human experts (super-human AI systems). This work is one of the very first steps towards the development of tools and methods that allow us to uncover hidden knowledge in highly capable AI systems, and empower human experts by helping them further improve their skills and understanding.

The super-human ability of AI systems may arise in a few different ways: pure computational power of machines, a new way of reasoning over existing knowledge, or super-human knowledge we do not yet possess. This work focuses on the last two cases. For simplicity, we refer to both as super-human knowledge from now on.

What does this mean from a research standpoint? The human representational space (H) has some overlap with the machine representational space (M) (see Figure 1 (Kim, 2022)). A representational space forms the basis of and gives rise to knowledge and abilities, which we are ultimately interested in. Thus, we use representational space and knowledge interchangeably – roughly speaking, H to represent what humans know and M to represent what a machine knows. There are things that both AI and humans know (M ∩H), things that only humans know (H −M), and things only machines know (M −H). Most existing research efforts only focus on (M ∩H), e.g., interpretability has tried to shoehorn M into (M ∩H), with limited success (Adebayo et al., 2018; Nie et al., 2018; Bilodeau et al., 2022). We believe that the knowledge gap represented by (M −H) holds the crucial key to empowering humans by identifying new concepts and new connections between existing concepts within highly performant AI systems. We already have evidence of cases when certain AI generations captivated the human imagination with ideas that were initially hard to grasp. One prominent example in the history of AI is the move 37 that AlphaGo made in a match with Lee Sedol. This move came as a complete surprise to the commentators and the player, and is still discussed to this day as an example of machine-unique knowledge. The vision to pursue super-human knowledge is ultimately for human-centered AI, and a world where human agency and capability do not come second. However, the question is–is this even possible?

This work is the first step towards discovering super-human knowledge and new connections of existing knowledge in (M −H). We focus on a domain that has inspired AI practitioners for decades, and captivated human imagination for centuries: the game of chess. Chess is an excellent playground to validate the existence and usefulness of set (M −H) for many reasons: chess knowledge has been developed over a long period of time, and the ground truth is much easier to validate compared to the frontiers of other fields, such as science or medicine. We also have a quantitative measure of the quality of play, both for human experts as well as machines, known as the Elo rating (Wikipedia contributors, 2023a).

Chess engines have performed at a super-human level for a long time, ever since DeepBlue’s match against Garry Kasparov. While early engines were based on human knowledge, the advent of AlphaZero (Silver et al., 2017) (AZ) showed a self-taught deep learning model achieve a superhuman capability in chess without any human knowledge. However, as humans, we have not yet been able to tap into their knowledge fully. Through analysis of AZ’s games, humans manually distilled patterns, such as its proclivity for playing on the flanks with moves like a4 or h4 (Sadler and Regan, 2019). However, this still analyses M through the lens of H, a bias that limits what we can find from M ∩H.

In this work, we aim to take the first step to change that by facilitating learning from the super-human knowledge in the (M −H) set of AZ. We hypothesise that (M −H) exists, and can be taught to humans.

Related work. Here, we review relevant prior work on concept discovery, interpretability of reinforcement learning systems, and the intersection of AI and chess.

In contrast to traditional feature or data-centric interpretability methods (Ribeiro et al., 2016; Lundberg and Lee, 2017; Sundararajan et al., 2017; Koh and Liang, 2017), concept-based methods use high-level abstraction, concepts, with the goal of providing model explanations to inform human practitioners (Bau et al., 2017; Kim et al., 2018; Alvarez-Melis and Jaakkola, 2018; Koh et al., 2020; Bai et al., 2022; Achtibat et al., 2022; Crabb ́e and van der Schaar, 2022). These types of explanations are shown to be useful in scientific and biomedical domains (Graziani et al., 2018; Sprague et al., 2019; Clough et al., 2019; Bouchacourt and Denoyer, 2019; Yeche et al., 2019; Sreedharan et al., 2020a; Schwalbe and Schels, 2020; Mincu et al., 2021; Jia et al., 2022), where experts’ concepts are highly relevant in decision making rather than individual low-level features.

Going beyond supervised concepts and probe datasets has also been investigated (Yeh et al., 2020; Ghorbani et al., 2019; Ghandeharioun et al., 2021) to discover concepts that a model represents without being limited to human labelled concepts. The concept is expressed using examples of training data (Yeh et al., 2020; Ghorbani et al., 2019) or by generating new data (Ghandeharioun et al., 2021). This work falls under methods to discover concepts but with a different goal of discovering and teaching humans new concepts rather than finding existing human concepts.

2.2 Generating explanations in Reinforcement Learning For trained RL systems, there is a pressing need for post-hoc RL interpretability methods. Input saliency maps (Wang et al., 2016; Selvaraju et al., 2019; Greydanus et al., 2018; Mundhenk et al., 2020) and tree-based models (Bastani et al., 2018; Roth et al., 2019; Coppens et al., 2019; Liu et al., 2019; Vasic et al., 2019; Madumal et al., 2020) have been a common approach. Saliencybased RL explainability approaches are not without issues, as they may suffer from unfalsifiability and be subject to cognitive bias (Atrey et al., 2019) as well as provably wrong results (Bilodeau et al., 2022). Visualizing the agent memory over trajectories (Jaunet et al., 2020) or extracting finite-state models (Koul et al., 2018) are explored to improve understanding of agents’ behavior, as well as leveraging Markov decision processes (Finkelstein et al., 2022; Zahavy et al., 2016) to generate explanations or detect sub-goals or emerging structures (Rupprecht et al., 2019).

Method. There are several possible definitions of a concept – varying from a human-understandable high-level feature to an abstract idea. In this work, we define concepts as a unit of knowledge. There are two key properties we focus on. The first is that a concept contains knowledge: information that is useful; in the context of machine learning, we take this to mean that it can be used to solve a task. For example, consider the concept of a beak. We can teach an algorithm or person (transfer of the knowledge) what a beak is. If the person grasps the beak concept, they can use it to identify birds. Second, a unit implies minimality; it is concise and irrelevant information has been removed.

There are many ways to operationalise this definition and properties, and we choose one of them: showing a concept can be transferred to another agent to help them solve a task (e.g., follow the strategy represented in a concept). Being able to do so implies that the concept is self-contained and useful for the task.

How do we represent concepts? We leverage rich literature that assumes concepts are linearly encoded in the latent space of a neural network (McGrath et al., 2022; Kim et al., 2018; Gurnee et al., 2023; Conneau et al., 2018; Tenney et al., 2019; Nanda, 2023). The latent space refers to the space spanned by post-activation features of a neural network. Although our assumption of linearity is a strong assumption, it has a surprising amount of empirical support: linear probing and related techniques have successfully extracted a wide range of complex concepts from neural networks across multiple domains (McGrath et al., 2022; Kim et al., 2018; Gurnee et al., 2023; Conneau et al., 2018; Tenney et al., 2019; Nanda, 2023). Although we may miss concepts with nonlinear representations, we nevertheless show that we can find useful concepts for our goal using purely linear representations.

What types of concepts do we aim to discover in the RL setting? We aim to discover concepts that give rise to a plan, where a plan is a deliberate sequence of actions optimizing for one or more relevant concepts. We take deliberate to mean that there is an underlying reason. More specifically, we assume a plan is motivated by one or more concepts. Although the terminal goal of a plan the same across states – maximizing the outcome (win or draw) – plans in a specific state will have more context-specific instrumental goals along the way, for instance, capturing a particular piece in an advantageous position, or maximising one’s board control. We assume that plans in similar contexts will share similar instrumental goals, and thus give rise to similar concepts.

Our method can be summarised into (1) excavating vectors that represent concepts in AZ using convex optimisation, (2) filtering the concepts based on teachability (whether it is transferable to another AI agent) and novelty (whether it contains some information that is not present in human games). The resulting set of concept vectors is then used to generate chess puzzles (chess positions and solutions), which are presented to human experts (top chess grandmasters) for final validation.

To find concepts, we develop a new method since (1) the model input is a mix of binary and realvalued inputs (e.g., saliency maps typically take as input continuous values and are generally not suitable for binary values) and (2) we want to develop an interpretability tool to analyse both parts of AZ machinery – the policy-value network and MCTS. Leveraging both the network and MCTS is crucial, since each component plays a different yet important role in deciding the move (see §8.3 for more detail). We formulate concept discovery as a convex optimisation problem. Using a convex optimisation framework is not new; many existing methods for finding concept vectors, such as non-negative matrix formulation, can often approximated as a convex optimisation problem (Ding et al., 2008).

Discussion. We investigate whether top chess grandmasters can successfully learn and subsequently apply the concepts we discovered (in §4). In particular, we investigate if it is possible to learn these concepts via exposure to a small set of prototypes of each concept. These prototypes are found using a concept vector as in §4.2.1, and then filtered based on a set of criteria aimed at identifying the most relevant high-quality prototypes (see §8.8 for more information). Learning from prototypes is similar to the established approaches to teaching chess. Chess students are often presented with puzzles sampled according to a theme (opening, chess position type, piece sacrifices, etc.), and practicing puzzle solving (i.e., finding the correct next moves) is one way of improving their overall strength and ability (Alisa Melekhina, 2014).

Henceforth, we use puzzles to denote prototypes and their ‘solutions’ (AZ’s selected move). Human evaluation with grandmasters follows three phases, similar to how teachability is measured §4.2.1:

• Phase 1: Measuring baseline performance. Each grandmaster provides solutions for a set of provided puzzles corresponding to a set of concepts. This phase determines the baseline performance: the number of puzzles in which the chess grandmaster gets the continuation correct before the learning phase.

• Phase 2: Learning from AZ’s calculations. The same puzzles as in Phase 1 are shown to chess grandmasters alongside the associated AZ’s suggested top line based on MCTS calculations for each puzzle. This serves as the simplest way of teaching.

• Phase 3: Measuring final performance. Grandmasters are tasked with providing solutions for a test set of unseen puzzles sampled from the same concepts they have seen in Phase 1. We Overall, we find that all study participants improve notably between phases 1 and 3, as shown in Table 4, suggesting that the chess grandmasters were able to learn and apply their understanding of the represented AZ chess concepts. The magnitude of improvement does not correlate with the chess player’s strength (i.e., Elo rating). Below, we discuss factors that may have influenced performance:

6.2 Qualitative analysis of the concepts and illustrative examples Overall, the grandmasters appreciated the concepts, describing them as ‘clever’ (Figure 8), ‘very interesting’ (Figure 16), and ‘very nice’ (Figure 18). Further, they found that the ideas often contained novel elements, commenting that the moves were ‘something new’ and even ‘not natural’ (Figures 16) and 10). Often, the grandmasters found the positions were very complex – making remarks such as that it was “very complicated - not easy to understand what to do”. Even when seeing AZ’s solutions, they remarked that it was a ‘very nice idea which is hard to spot’ (Figure 22).

6.2.3 Differences between humans and AZ The qualitative examples suggest that AZ has different priors over the relevance of concepts in a chess position than humans. Human chess players formulate and adopt heuristic chess principles to inform their analysis, predisposing them to biases that influence which concepts they deem relevant for specific chess positions. An example is the three ‘golden rules’ of the opening: control the centre, develop your pieces, and bring your king to safety (Hansen, 2021; Brunia and van Wijgerden, 2021; King, 2000). Consequentially, in opening, humans may focus on moves that align with these guidelines. Instead, AZ is self-taught and does not seem to have the same priors over chess concepts as humans. We believe this lack of prior allows AZ to be more flexible – it can apply concepts to various different chess positions and change plans quickly.

Conclusion. Our research represents a first step toward understanding the potential of human learning from Artificial Intelligence (AI). In this work, we focused on AlphaZero (AZ) – an AI model that learned to play chess at a super-human level through self-play without prior knowledge or human bias. Through spectral analysis, we show that AZ’s games encode features that are not present in human games, providing evidence for the existence of super-human knowledge. To extract knowledge from AZ’s representational space, we developed a framework to uncover new chess concepts in an unsupervised fashion. We discover unsupervised concepts by leveraging AZ’s training history to curate a set of complex chess positions. We ensured each concept was informative, by verifying that the concept can be taught to another AI agent, and novel, through a spectral analysis of human and AZ games. Communicating novel concepts requires a common language between humans and AI. We bypass the need for this language by creating puzzles for each concept.

We collaborated with four world-top grandmasters to (1) validate the human capacity to comprehend and apply these concepts by studying AZ’s concept prototypes and (2) improve our understanding of the differences between AZ’s and humans’ chess representation space. All four grandmasters improved their performance after learning concepts compared to baseline performance. We speculate that the differences between AZ and humans may stem from (1) prior biases over concepts, including their perceived applicability, importance, and how they can be combined with other concepts. For example, AZ shows a reduced emphasis on factors such as material value and is more agile in switching between playing on different sides of the board. (2) a difference in the motivation and objectives when playing chess; AZ is trained to accurately evaluate the current chess position, There are several aspects of the work that could be further explored. In our work, we found a subset of all possible concepts. For example, we limited our investigation to linear sparse concept vectors. However, other concepts may be discovered in the form of non-linear vectors.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do AI systems determine and balance multiple competing objectives? How do users confuse explanation quality with actual system accuracy? Can mechanistic interpretability methods reliably reveal what models actually know? How do neural networks learn compositional structure from training? When do simpler collaborative filtering approaches outperform complex LLM recommenders? How does diversity prevent model convergence on superficial patterns? How do sequence length and task type interact with sparsity tolerance? Why do vector embeddings fail at capturing task-relevant relationships? Does pretraining establish the ceiling for what reward learning can improve? Can minimal training unlock latent reasoning already present in base models?