Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Paper · arXiv 2608.12036 · Published August 12, 2026
Deep Research Agents

AI models are increasingly used in scientific discovery and human decision-making. Yet how AI models work and what risks they pose remain poorly understood. As AI development becomes faster and more automated, research on the mechanisms underlying AI remains largely manual, widening the gap between model capabilities and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI. To ground novel mechanism hypotheses, we construct a scientific knowledge graph of 13,000 studies on AI mechanisms, alongside a multidisciplinary database of 43 million papers spanning 26 fields. For reliable experiment execution, we curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates higher-quality mechanism hypotheses and executes experiments more reliably. Across four case studies, Mechanist autonomously discovers new model behaviors and their underlying mechanisms, and translates these discoveries into mechanism-guided interventions and interdisciplinary design.

Introduction. Artificial Intelligence (AI) models are rapidly evolving from tools for daily chat into intelligent systems that guide problem formulation, experimental design, and human decision-making [1, 2, 3]. The growing success of AI models raises a fundamental question [4, 5, 6]: how do these AI models acquire knowledge about the world, form beliefs 7, reason, and act? It also remains unclear whether AI models harbour latent risks or exhibit unreliable behaviors in critical domains such as healthcare, finance, and chemical manufacturing [9, 10, 11]. More critically, AI development is accelerating and becoming increasingly automated, outpacing progress in understanding and controlling the mechanisms underlying AI[1, 4]. Discovering the mechanisms of model intelligence can reveal how models operate internally, elucidate the principles underlying their behavior, identify risks early, and enable adaptive control over and targeted improvements to model behavior [9, 10, 11]. Despite its importance, mechanistic understanding of AI remains difficult to obtain and hard to scale [10, 12, 13].

Discussion / Conclusion. We introduce Mechanist, a scientific instrument for the autonomous investigation of mechanisms underlying AI models. To support Mechanist, we design a large-scale knowledge graph of interpretability and a library of 32 foundational methods of mechanistic analysis. These resources enable Mechanist to formulate high-quality hypotheses, execute experiments reliably, establish robust causal evidence, and iteratively refine mechanistic explanations. Mechanist can uncover previously unrecognized risks in models deployed in scientific laboratory settings, reveal how AI models represent world knowledge, and improve model capabilities across computer science and other scientific domains. Then, we discuss related topics in the following section. AI for science.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How faithfully do LLMs reflect their actual reasoning in outputs and explanations? Why do reasoning models fail at systematic problem-solving and search? Can AI-generated outputs constitute genuine knowledge or valid claims? What limits mechanistic interpretability's ability to characterize models? How can AI agents autonomously learn and transfer skills across tasks? How should models express uncertainty rather than forced confident answers? How do LLMs distinguish causal reasoning from temporal and semantic associations? How should we design LLM systems to maintain alignment and control? How can identical external performance mask different internal representations? Is model self-awareness based on genuine introspection or pattern matching? Why do continual learning scenarios trigger catastrophic forgetting and interference? Do reasoning traces faithfully represent or merely mimic actual model reasoning? Do language model representations contain causally steerable task-specific features? How do neural networks separate factual knowledge from reasoning abilities? Why does supervised fine-tuning improve accuracy while degrading reasoning quality? Do language models develop causal world models or rely on statistical patterns?