Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
AI models are increasingly used in scientific discovery and human decision-making. Yet how AI models work and what risks they pose remain poorly understood. As AI development becomes faster and more automated, research on the mechanisms underlying AI remains largely manual, widening the gap between model capabilities and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI. To ground novel mechanism hypotheses, we construct a scientific knowledge graph of 13,000 studies on AI mechanisms, alongside a multidisciplinary database of 43 million papers spanning 26 fields. For reliable experiment execution, we curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates higher-quality mechanism hypotheses and executes experiments more reliably. Across four case studies, Mechanist autonomously discovers new model behaviors and their underlying mechanisms, and translates these discoveries into mechanism-guided interventions and interdisciplinary design.
Introduction. Artificial Intelligence (AI) models are rapidly evolving from tools for daily chat into intelligent systems that guide problem formulation, experimental design, and human decision-making [1, 2, 3]. The growing success of AI models raises a fundamental question [4, 5, 6]: how do these AI models acquire knowledge about the world, form beliefs 7, reason, and act? It also remains unclear whether AI models harbour latent risks or exhibit unreliable behaviors in critical domains such as healthcare, finance, and chemical manufacturing [9, 10, 11]. More critically, AI development is accelerating and becoming increasingly automated, outpacing progress in understanding and controlling the mechanisms underlying AI[1, 4]. Discovering the mechanisms of model intelligence can reveal how models operate internally, elucidate the principles underlying their behavior, identify risks early, and enable adaptive control over and targeted improvements to model behavior [9, 10, 11]. Despite its importance, mechanistic understanding of AI remains difficult to obtain and hard to scale [10, 12, 13].
Discussion / Conclusion. We introduce Mechanist, a scientific instrument for the autonomous investigation of mechanisms underlying AI models. To support Mechanist, we design a large-scale knowledge graph of interpretability and a library of 32 foundational methods of mechanistic analysis. These resources enable Mechanist to formulate high-quality hypotheses, execute experiments reliably, establish robust causal evidence, and iteratively refine mechanistic explanations. Mechanist can uncover previously unrecognized risks in models deployed in scientific laboratory settings, reveal how AI models represent world knowledge, and improve model capabilities across computer science and other scientific domains. Then, we discuss related topics in the following section. AI for science.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How faithfully do LLMs reflect their actual reasoning in outputs and explanations?- What distinguishes LLM fabrication from genuine theoretical reasoning?
- What internal mechanisms explain LLM reasoning and representation limits?
- Can mechanistic interpretability reveal how ideologies decompose into simpler features?
- How much do mechanistic interpretability findings reflect true reasoning architecture?
- Are detection and identification of injections truly separable in neural circuits?
- Do causal rules enforce robustness that statistical patterns alone cannot maintain?
- What makes causal belief networks more auditable than prompted personas?
- Can causal models be extended to include non-causal cognition?