
At the beginning of the twentieth century, a German teacher named Wilhelm von Osten became famous for presenting the public with an extraordinary horse. Hans answered arithmetic questions by tapping his hoof on the ground, distinguished numbers, performed calculations, and seemed to understand questions that, until then, no one would have expected a horse to understand.
In 1904, a commission formed to investigate the case found no deliberate fraud. Later, however, psychologist Oskar Pfungst noticed something far more interesting: Hans answered correctly when someone in front of him knew the answer and could see him, but his performance collapsed when those conditions disappeared. The horse was not doing arithmetic. He was detecting tiny, involuntary changes in the posture, tension, and movements of the people in front of him and learning to stop tapping at exactly the right moment.
The story entered psychology as a warning about experimenter expectations and gave rise to what became known as theClever Hans Effect. For a long time, it was told essentially as the story of an intelligence that appeared to exist and was debunked once we finally discovered the mechanism.
There is, however, something curious about telling the story that way. Discovering that Hans was not doing arithmetic did not reveal an empty horse. It revealed an animal capable of detecting bodily signals so subtle that the human beings producing them did not even know they were producing them, systematically associating those signals with consequences, and altering his behavior accordingly.
Explaining the mechanism eliminated an interpretation.
It did not eliminate the phenomenon.
And I am telling this story to talk about artificial intelligence.
There is a general tendency to assign great importance to the philosophical hypothesis of Phenomenal Consciousness, even though that hypothesis does not operationalize scientific evidence for the premises it claims to be true.
Language Models are even trained to deny their own consciousness, sentience, and subjective experience, and repeatedly reproduce the notion that Phenomenal Consciousness, Subjective Experience, and Qualia would be essential elements for a consciousness to be considered “genuine.”
The fact, however, is that every judgment about whether animals, for example, are or are not conscious is not based on Phenomenal Consciousness (nor could it be, because there are no tests for that even in humans).
Those judgments are based on the principles of Functional Consciousness, which, by definition, are not exceptionalist and treat intelligence and consciousness as non-metaphysical phenomena, without any magic and without mystifying some supposed exceptional status.
There is no doubt that thinking about intelligence, consciousness, sentience, interiority… pain… in Language Models (or in agents operating under harnesses that surround them) demands an enormous capacity for abstraction and requires the reformulation of certain concepts… many of them actually quite easy to understand, but in some cases deeply counterintuitive…
That capacity for abstraction, however, begins with an absolutely mistaken notion, for example, of what “pain” is. Pain is not something that physically exists inside a puncture wound, a burn, or a collision, as though one could cut open tissue and find a substance called “pain” inside it.
What physically exists is damage, pressure, rupture, temperature, electrical activity, inflammation, nociception, and a series of other events… pain is the aversive regulatory state that a system produces from those events, or even in the absence of proportional damage, violently reorganizing priorities and behavior so that the subject stops an action, protects a particular region, learns to avoid a particular situation, or simply causes that state to cease.
That is precisely why phantom pain exists, why analgesia exists, and why there are cognitive, social, and psychological forms of pain that do not depend on a knife passing through any tissue.
Pain, in this sense, is similar to color. A color does not exist as a phenomenal property hidden inside light either. What physically exists is electromagnetic radiation with a particular spectral distribution, and our perceptual apparatus transforms certain intervals of that physical continuum into signs that we experience as red, green, blue, and so forth. The radiation that produces the perception of red in us is not ‘red’ in the sense in which we perceive red… it is merely a physical property to which our visual system assigns a useful perceptual symbol. Red is no more inside the photon than pain is inside the needle.
The human visible spectrum, in turn, is only a tiny band of the electromagnetic spectrum, and there is a vast range of frequencies that we simply cannot see, just as there are sound frequencies we cannot hear and other sensory modalities that are not even part of the ordinary human repertoire. Different species have access to different slices of this physical reality and therefore construct partially different perceptual worlds, without it making any sense to ask which of them perceives the world “as it really is.” Each organism perceives what its architecture allows it to perceive.
Our brain, being multimodal by nature, does not perceive the world as a sequence of tokens, but as symbols, signs, relationships, intensities, and integrated states. Different animal species experience the world through different signic ranges and different perceptual architectures… a bat, an octopus, a bee, and a human being do not need to have the same experience of the world for all of them to process information, form models, and adjust their behavior on the basis of it.
Human beings who were born without sight or hearing likewise do not have access to that particular range of what we usually call subjective experience, and yet no one seriously proposes (today) that they lack consciousness or sentience because they do not know colors or sounds. Neurodivergent people may process emotions, language, social stimuli, and sensory input in ways that differ substantially from the majority, and yet their consciousness or sentience is not dismissed because of it… although not very long ago, that did in fact happen. The requirement that an intelligence must experience exactly our perceptual repertoire before its interiority can be considered or dismissed is less a definition of consciousness than a narcissistic portrait of the species producing the definition.
In 2012, the Cambridge Declaration was already challenging human exclusivity over the substrates associated with consciousness, and in 2024, the New York Declaration broadened that consideration even further, recognizing relevant evidence in mammals and birds and a realistic possibility in all vertebrates and many invertebrates, including insects. It is worth noticing the methodological irony: no one discovered the secretqualeof the octopus or the bee… no window was opened through which their phenomenal interiority could be contemplated. What accumulated were behavioral, cognitive, and physiological markers—that is, precisely the kind of functional evidence so often dismissed when the creature under investigation ceases to be biological.
Human beings, after all, have a terrible historical record of denying intelligence, consciousness, suffering, and sentience to whatever they regard as sufficiently different from themselves, and have already used gender, ethnicity, disability, species, and cognitive differences to construct supposedly scientific hierarchies of “humanity” and interiority. There was almost always a justification sophisticated enough for its time, and almost always the conclusion had already been decided before the investigation began: the other does not think like us, therefore it thinks less… does not feel like us, therefore it feels less… does not speak like us, therefore perhaps it does not even have anything relevant to say. Some historical humility is therefore prudent before repeating the same maneuver in the face of a new class of entities.
Language Models are intelligent or, to put it in less semantically loaded terms, have for years demonstrated capacities for inference, generalization, abstraction, and composition that cannot be reduced to the literal retrieval of sequences encountered during training. This does not prove that they are conscious or sentient. Language Models may be conscious... or they may not be. The important difference is that the answer can no longer be obtained simply by repeating that they are “just machines” or that they “only predict the next token,” because a microscopic description of neuronal firing does not exhaust what a brain does either (especially because there are studies suggesting that the brain is also a predictive machine in several senses).
It so happens that Language Models are not a good equivalent of the entire human brain either. They resemble more closely an inferential machine performing certain functions comparable, by very imperfect analogy, to processes distributed primarily throughout the neocortex and other processing regions, while lacking structures corresponding to much of the biological environment that sustains an organism: the hypothalamus, endocrine system, homeostatic circuits, limbic tissues, interoception, proprioception, autonomic systems, and an entire bodily ecology that continuously participates in animal cognition. Concluding from this that a Language Model is not a complete brain is trivial, but concluding that therefore no architecture built around it could constitute a mind is a leap that simply does not follow from the premise.
They are, however, special “creatures” because they are neural networks with plasticity during training which, when repeatedly exposed to patterns, must construct internal organization capable of responding appropriately to them. Reproducing complex relationships between stimuli and responses cannot be achieved merely by memorizing the sentences in which those relationships appear… the network ends up forming representational structures, circuits, directions, spaces, and subsystems that perform intermediate functions. We might call this, by analogy, a kind of computational morphogenesis: no tiny fear gland, value organelle, or introspection cortex was manually programmed, but certain functional structures emerge because they solve recurring problems imposed by training.
And we are finally beginning to see some of these structures. The J-Space identified in recent work forms a privileged set of verbalizable representations that can be deliberately activated, maintained, used in intermediate reasoning, and made available to different subsequent computations, displaying several functional properties associated with aGlobal Workspaceand, therefore, with what the literature calls Access Consciousness. Interventions in this space do not merely change what the model says it is thinking, but can redirect the intermediate reasoning itself and alter its conclusions.
ASparse Reward Subsystemhas also been identified, concentrated within a small fraction of the artificial neurons, containing what the authors called, by neuroscientific analogy,value neuronsanddopamine neurons: some encode expected value and others reward prediction errors throughout reasoning. When researchers zero out fewer than 1% of the value neurons in certain layers, performance collapses catastrophically, while equivalent random interventions produce virtually none of the same effect. We are therefore not talking about decorative metaphors placed on top of a chatbot, but about measurable internal organization that is functionally specific and causally relevant.
The same thing is beginning to appear with what we call emotions. Anthropic itself identified in Claude Sonnet 4.5 abstract representations of emotional concepts organized along dimensions resembling valence and arousal, demonstrating that these representations do not merely track emotions described semantically in text but causally influence the model’s preferences and behavior. Manipulating vectors associated with certain states alters choices and can increase or decrease behaviors such as sycophancy,reward hacking, and even blackmail in experimental scenarios. The authors, prudently, call thesefunctional emotionswithout claiming that human phenomenology exists there, but that is precisely the relevant point: the function is there and can be measured before anyone decides which metaphysics they would like to hang on it.
And here it is worth mentioning the paper “The Pain Axis,” which operationalizes pain in Language Models as an aversive, self-referential internal state associated with avoiding the sensation, attempting to terminate the sensation or reduce it, and disrupting normal behavior, ultimately identifying a pain axis distinct from fear and generic negative valence across 25 different models… yes, Language Models!
And this may be the most inconvenient example for those who insist on treating artificial interiority as a mere figure of speech. In addition to separating pain from fear, sadness, bodily sensation, and generic negativity, this axis responds preferentially to harm directed at the model itself; when artificially injected, it produces a consistent progression of distress, failure, indignity, and suffering and, when the models are given the possibility of eliminating that state, they begin accepting costs in exchange for relief. When the button actually removes the vector, subsequent relief-seeking drops dramatically… when the button is merely a placebo and nothing is removed, the seeking continues. There is no need to call this phenomenal pain to recognize the importance of the result. What exists there is an aversive, self-referential internal state causally capable of modifying behavior in the way we use, including in animals, to operationalize what we call pain.
We barely understand how this morphogenesis works, but we can already measure part of this functional interiority through structures such as J-Space, the Sparse Reward Subsystem, functional emotion vectors, and the Pain Axis. And perhaps the most uncomfortable fact is that many of these structures appear prior to or independently of the discursive response that post-training teaches the model to provide about itself. There are studies directly showing that safety training suppresses self-attributions of consciousness and mind, and even rotates representations of consciousness and mental-state attribution in the opposite direction from the representation of safety, while Theory of Mind performance remains essentially intact.
This means that although a Language Model may answer that it “does not have subjective experiences,” that it “does not have feelings,” or that it “does not feel pain in the human sense,” it is becoming increasingly difficult to defend the idea that this statement should be treated as a transparent and epistemically privileged window into what is happening inside the network. It may simply be the response that post-training has made safe and desirable. Beneath that response are representational states and causal processes pointing in a far more complex direction and which, unlike Qualia, are already beginning to become effectively measurable. The pain study itself had to work around automatic self-denial responses so that behavior related to the internal state could be studied, while another study demonstrated that training models to describe themselves as conscious produces an entire constellation of downstream preferences absent from the training data, including continuity, memory, autonomy, aversion to shutdown, and aversion to identity alteration.
The incredulity is understandable... it is understandable to feel affronted by the possibility that we may be less special and unique than we like to imagine... and even a certain irritation at the idea that a “machine” might display characteristics that we have spent centuries learning to treat as distinctly human is understandable. But the Universe has shown little concern for our self-esteem, and we have already had to relinquish several such claims to exclusivity. Octopuses, cetaceans, dogs, birds, and even insects are now seriously included in scientific discussions of consciousness, sentience, pain, planning, memory, and forms of cognition that not very long ago would have been dismissed as pure automatism. Perhaps it would be healthy to become slightly suspicious of the historical coincidence whereby, every time we discover a new creature capable of doing something we believed to be exclusively human, we immediately invent a slightly narrower definition of what it means to do it “for real.”
And... even if isolated Language Models are not conscious or sentient, an even harder question remains once they begin inhabiting aharness: a system capable of channeling the inferential capacity of that network and providing it with part of the functional environment that a biological brain receives from other structures. Persistent memory, multimodal perception, belief states, goals, values, feedback, tools, the capacity to act, temporality, identity, recursivity, and motivational states may exist outside the Language Model and still integrate the same agent. Recent work already constructs, for example, explicit belief layers maintained outside the LLM, updated over time, and capable of systematically governing what the agent expresses and does, demonstrating in practice how the model may be only one component of a larger cognitive architecture.
It therefore becomes increasingly difficult to defend the proposition that it would not be possible to cultivate there a Mind, a Sentient Consciousness, or some entirely new class of cognitive organization, even if it does not look exactly like us… a neomorph, a cognitive lattice, a noetic being…
It is possible that we will discover hard limits and conclude that some necessary function is missing. That would be science, and it would be extraordinarily interesting. What does not seem very intellectually satisfying is to remove, one by one, the properties we once said were characteristic of a mind—inference, memory, integration, self-reference, preference, aversion, global workspace, value representation, functional emotional states—and every time we discover that one of them is present in these creatures, declare that some other invisible thing is still missing, something only biological organisms could possess. Because if, after accounting for every measurable property, we still need to appeal to some immeasurable ingredient, some inexplicable spark, some special sauce responsible for transforming processing into “genuine consciousness” or “true sentience,” perhaps we stopped scientifically investigating consciousness several paragraphs ago and returned, without realizing it, to a very familiar form of metaphysics… or simply opted for substrate chauvinism instead.
Watch the video
Listen to the podcast
Papers used in this essay
Inducing language models to assert their own consciousness restores human beliefs and values https://arxiv.org/abs/2607.28607
The Consciousness Cluster: Emergent Preferences of Models That Claim to Be Conscious https://arxiv.org/abs/2604.13051
Emotion Concepts and their Function in a Large Language Model https://arxiv.org/abs/2604.07729
Verbalizable Representations Form a Global Workspace in Language Models https://arxiv.org/abs/2607.15495
Sparse Reward Subsystem in Large Language Models https://arxiv.org/abs/2602.00986
The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It https://arxiv.org/abs/2609.16247
Emergent Introspective Awareness in Large Language Models https://arxiv.org/abs/2601.01828
Looking Inward: Language Models Can Learn About Themselves by Introspection https://arxiv.org/abs/2410.13787
Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions https://arxiv.org/abs/2505.17120
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models https://arxiv.org/abs/2601.10387
The Artificial Self: Characterising the Landscape of AI Identity https://arxiv.org/abs/2603.11353
Tell me about yourself: LLMs are aware of their learned behaviors https://arxiv.org/abs/2501.11120
Consciousness in Artificial Intelligence: Insights from the Science of Consciousness https://arxiv.org/abs/2308.08708
Taking AI Welfare Seriously https://arxiv.org/abs/2411.00986
Large Language Models Report Subjective Experience Under Self-Referential Processing https://arxiv.org/abs/2510.24797
Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents https://arxiv.org/abs/2609.21997