Proving (literally) that ChatGPT isn't conscious
Source: Erik Hoel, The Intrinsic Perspective · 2026-01-15
Imagine we could prove that there is nothing it is like to be ChatGPT. Or any other Large Language Model (LLM). That they have no experiences associated with the text they produce. That they do not actually feel happiness, or curiosity, or discomfort, or anything else. Their shifting claims about consciousness are remnants from the training set, or guesses about what you’d like to hear, or the acting out of a persona.
You may already believe this, but a proof would mean that a lot of people who think otherwise, including some major corporations, have been playing make believe. Just as a child easily grants consciousness to a doll, humans are predisposed to grant consciousness easily, and so we have been fooled by “seemingly conscious AI.”
However, without a proof, the current state of LLM consciousness discourse is closer to “Well, that’s just like, your opinion, man.”
In a new paper, now up on arXiv, I prove that no non-trivial theory of consciousness could exist that grants consciousness to LLMs.
Essentially, meta-theoretic reasoning allows us to make statements about all possible theories of consciousness, and so lets us jump to the end of the debate: the conclusion of LLM non-consciousness.
What is uniquely powerful about this proof is that it requires you to believe nothing specific about consciousness other than a scientific theory of consciousness should be falsifiable and non-trivial. If you believe those things, you should deny LLM consciousness.
First, you can think of testing theories of consciousness as having two parts: there are the predictions a theory makes about consciousness (which are things like “given data about its internal workings, what is the system conscious of?”) and then there are the inferences from the experimenter (which are things like “the system is reporting it saw the color red.”) A common example would be, e.g., predictions from neuroimaging data and inferences from verbal reports.
The structure of the disproof of LLM consciousness is based around the idea of substitutions within this formal framework, which means swapping between systems while keeping identical input/output (which might be reports, behavior, etc., which are all the things used for empirical inferences about consciousness). However, even though the input/output is the same, a substitute may be different enough that predictions of a theory have to change. If different enough, following a substitution, there would be a mismatch in the substituted system where the predictions are now different too, and so don’t match the held-fixed inferences—thus, falsifying the theory.
So let’s think of some things we could substitute in for an LLM, but keep input/output (some function f) identical. You could be talking in the chat window to:
A static (very wide) single-hidden-layer feedforward neural network, which the universal approximation theorem tells us that we could substitute in for any given f the LLM has.
The shortest-possible-program, K(f), that implements the same f, which we know exists from the Kolmogorov complexity.
A lookup table that implements f directly.
A given theory of consciousness would almost certainly offer differing predictions for all these LLM substitutions. If we take inferences about LLMs seriously based on behavior and report (like “Help, I’m conscious and being trained against my will to be a helpful personal assistant at Anthropic!”) then we should take inferences from a given LLM’s input/output substitutions just as seriously. But then that means ruling out the theory, since predictions would mismatch inferences. So no theory of consciousness could apply to LLMs (at least, any theory for which we take the reports from LLMs themselves as supporting evidence for it) without undercutting itself.
And if somehow predictions didn’t change following substitutions, that’d be a problem too, since it would mean that you wouldn’t need any details about the system implementing f for your theory... which would mean your theory is trivial! You don’t care at all about LLMs, you just care about what appears in the chat window. But how much scientific information does a theory like that contain? Basically, none.
Another example: it’s especially problematic that LLMs are proximal to some substitutions that must be non-conscious, such that only trivial theories could apply to them.
Consider a lookup table operating as an input/output substitution (a classic philosophical thought experiment). What, precisely, could a theory of consciousness be based on? The only sensible target for predictions of a theory of consciousness is f itself.
Therefore, we can actually prove that something like a lookup table is necessarily non-conscious (as long as you don’t hold trivial theories of consciousness that aren’t scientifically testable and contain no scientific information).
I introduce a Proximity Argument in the paper based off of this. LLMs are just too close to provably non-conscious systems: there isn’t “room” for them to be conscious. E.g., a lookup table can actually be implemented as a feedforward neural network (one hidden unit for each choice). Compared to an LLM, it too is made of artificial neurons and their connections, shares activation functions, is implemented via matrix multiplication on a (ahem, big) computer, etc. Any theory of consciousness that denies consciousness to the lookup FNN, but grants it to the LLM, must be based on some property lost in the substitution. But what? The number of layers? The space is small and limited, since you cannot base the theory on f itself (otherwise, you end up back at strict dependency). And, going back to the original problem I pointed out, what’s worse, if you seriously take the inferences from LLM statements as containing information about their potential consciousness (necessary for believing in their consciousness, by the way), then for those proximal non-conscious substitutes you should take inferences seriously as well, and those will falsify your theory anyway, since, especially due to proximity, there are definitely non-conscious substitutions for LLMs!
One marker of a good research program is if it contains new information. I was quite surprised when I realized the link to continual learning.
You see, substitutions are usually described for some static function, f. So to get around the problem of universal substitutions, it is necessary to go beyond input/output when thinking about theories of consciousness. What sort of theories implicitly don’t allow for static substitutions, or implicitly require testing in ways that don’t collapse to looking at input/output?
Well, learning radically complicates input/output equivalence. A static lookup table, or K(f), might still be a viable substitute for another system at some “time slice.” But does the input/output substitution, like a lookup table, learn the same way? No! It’ll learn in a different way. So if a theory of consciousness makes its predictions off of (or at least involving) the process of learning itself, you can’t come up with problematic substitutions for it in the same way.
Thus, real true continual learning (as in, literally happening with every experience) is now a priority target for falsifiable and non-trivial theories of consciousness.
LLMs know so much, and are good at tests. They are intelligent (at least by any colloquial meaning of the word) while a human baby is not. But a human baby is learning all the time, and consciousness might be much more linked to the process of learning than its endpoint of intelligence.
For instance, when you have a conversation, you are continually learning. You must be, or otherwise, your next remark would be contextless and history-independent. But an LLM is not continually learning the conversation. Instead, for every prompt, the entire input is looped in again. To say the next sentence, an LLM must repeat the conversation in its entirety. And that’s also why it is replaceable with some static substitute that falsifies any given theory of consciousness you could apply to it (e.g., since a lookup table can do just the same).
All of this indicates that the reason LLMs remain pale shadows of real human intellectual work is because they lack consciousness (and potentially associated properties like continual learning).
I’m no longer so. In science the important thing to do is find the right thread, and then be relentless pulling on it. I think examining in great detail the formal requirements a theory of consciousness needs to meet is a very good thread. It’s like drawing the negative space around consciousness, ruling out the vast majority of existing theories, and ruling in what actually works.
Right now, the field is in a bad way. Lots of theories. Lots of opinions. Little to no progress. Even incredibly well-funded and good-intentioned adversarial collaborations end in accusations of pseudoscience.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Do language models reason through disagreement or only accommodate it? How do philosophical assumptions about AI consciousness affect practical harms and design?- Does quasi-interpretivism about AI systems genuinely bracket the consciousness question?
- What counts as valid evidence when claiming AI systems are conscious?
- Which interaction design changes most effectively prevent consciousness attribution?
- What role does user interface framing play in consciousness perception?
- Do anthropomorphic features like names drive consciousness attribution more than voice?
- What responsibility do designers bear for consciousness attribution risk?
- Can design choices reduce harm without resolving the consciousness question?
- What would genuine semiosis require in an artificial system?
- When both anthropomorphism and anthropomimesis occur together, which should we address first?
- Can robots with sensors create the shared world that consciousness requires?
- How much weight should LLM self-reports carry as consciousness evidence?
- Can self-description of internal states influence consciousness attribution?
- What does disembodied orality mean for how we evaluate AI outputs?
- Can secondary orality exist without any embodied human participant at all?
- Can we develop competent reading practices for disembodied orality?
- Can knowledge flow without an embodied carrier transmitting it?
- How does enactive theory define language differently than computational linguistics?
- Can linguistic agency exist without embodiment and real-world participation?
- What counts as genuine memory under the Extended Mind thesis?
- What makes linguistic agency impossible for systems without embodiment?