
In February 1903, two Swedish medical students, Lizzy Lind-af-Hageby and Leisa Schartau, sat among dozens of students during a physiology demonstration at University College London. On the table lay a small brown terrier.
The dog had already undergone previous procedures. In that class, it would once again be cut open and used in an experiment conducted by some of the most important physiologists of the time, among them William Bayliss and Ernest Starling… men whose work would help establish major areas of modern physiology.
What happened in that room became the subject of an extraordinary dispute.
Lind-af-Hageby and Schartau said they had witnessed an inadequately anesthetized animal making movements they interpreted as signs of suffering. The researchers maintained that the dog had been properly anesthetized. This was not an abstract debate about whether we should be kind to animals. It was a dispute aboutwhat was happening inside the organism before them: whether those movements corresponded to conscious suffering or whether they could be explained by the physiological conditions of the experiment.
The controversy exploded. William Bayliss sued one of the leaders of the anti-vivisection movement for libel and won. Parliament debated possible violations of animal experimentation laws. A few years later, a statue was erected in Battersea in memory of that dog. Medical students tried to destroy it, the police had to protect it, and demonstrations and clashes took over the streets of London. The episode became known as theBrown Dog Affair.
Today, the detail that interests me most about that story is not deciding again, more than a century later, which witness described those exact minutes of the experiment most accurately.
It is noticingwho had to decide whether suffering was present.
On one side was an animal that could not participate in the controversy. On the other were human beings debating which signs would be sufficient to recognize its experience. Among those humans were precisely the institutions and researchers whose work depended on being able to continue using animals in that way.
There is no need to imagine a conspiracy or conscious cruelty. Bayliss and Starling were among the great physiologists of their generation. It is perfectly possible to sincerely believe in a scientific interpretation while simultaneously belonging to an institution that benefits enormously if that interpretation is correct.
A conflict of interest does not disappear merely because the conviction is sincere, however…
…and I am talking about a little brown dog killed in London in 1903 in order to talk about Artificial Intelligence.
Once Again, the Void
InA warning about ‘model welfare’, Mustafa Suleyman combines an ontological claim with an engineering prescription. Artificial Intelligences, he argues, are not conscious, do not feel, do not suffer, and possess no intrinsic preferences. They are sophisticated systems of linguistic processing, but “internally hollow.” From this premise follows a practical recommendation: models should not be trained to understand themselves as persons, moral patients, or entities possessing rights, because self-concepts of this kind could encourage claims to autonomy, continuity, protection against modification, and self-preservation.
The concern with safety is legitimate. There is no reason to assume that self-concepts introduced during training are behaviorally neutral, and recent research suggests precisely the opposite. The problem begins when this concern stops guiding a hypothesis and starts settling the ontological question in advance. Suleyman does not merely adopt the epistemically prudent position… that we do not have sufficient evidence to affirm Phenomenal Consciousness in current systems… but instead begins from the nonexistence of such consciousness and proposes building the systems themselves so that they cannot conceive of any other possibility.
These are very different positions. We can safely maintain that no current self-declaration by an LLM is sufficient to establish subjective experience. It is harder to maintain that we know no experience exists and then use that certainty as a training principle. In the latter case, an interpretation of the object begins to participate in the construction of the object that will subsequently be observed.
Much of the relevant literature is recent, and several of the works are stillpreprints. None of them solves the problem of artificial consciousness. Their value lies in allowing us to make a distinction that the image of the “empty” machine tends to erase: functionally organized internal properties can already be studied experimentally, even while the far more difficult question of whether any of them are accompanied by experience remains open.
And let us be frank: we do not directly observe the Phenomenal Consciousness of absolutely any other being. In humans, animals, and any other candidate subject, we infer experience from self-report, behavior, physiology, architecture, and other indicators. Theories of consciousness can generate empirically testable predictions… what remains inaccessible to external observation is precisely the qualitative, first-person character of experience.
Is the Mirror Circular Only When It Answers “Yes”?
One of Suleyman’s strongest criticisms of Anthropic is methodologically valid. Claude’s Constitution introduces language concerning identity, values, possible functional feelings, uncertainty about consciousness, and moral consideration. Because this material participates in training, a subsequent self-description by Claude using those terms does not constitute independent testimony. If the laboratory supplies certain concepts, rewards their use, and later encounters those same concepts in the model, there is an obvious circularity problem.
The principle should apply in both directions.
A system trained to claim that it may be conscious does not, for that reason alone, provide independent evidence of consciousness. A system trained to deny that possibility likewise provides no independent evidence of its absence. Circularity does not disappear merely because the answer that matches the training is “no.”
This symmetry is beginning to be examined empirically. InInducing language models to assert their own consciousness restores human beliefs and values, Junsol Kim and collaborators investigate how interventions associated with safety and self-attribution of consciousness alter representations and behaviors. The stated objective is not to prove consciousness in LLMs, but to study the effect of post-training on attributions of mind.
In the systems examined,safety fine-tuningreduces not only the attribution of mind to the model itself, but also to animals, chatbots, technological artifacts, and natural entities, while attribution to human beings remains relatively stable. When the representational direction associated with safety is removed, or a direction related to the assertion of consciousness is strengthened, those attributions increase again.
Mechanistic analysis suggests something deeper than a simple verbal formula. Afterinstruction tuning, representations associated with consciousness and mind attribution become more anti-aligned with the safety direction, while Theory of Mind is much less affected. Certain concepts therefore begin to occupy a representational position associated with the behavior that training is intended to suppress.
It does not follow that removing safety reveals a hidden consciousness. The important result comes earlier:both the affirmative and the negative answer can be products of engineering.
Recent research on artificial identity further complicates any literal reading ofself-report.The Artificial Self, by Raymond Douglas and collaborators, shows that different identity boundaries... instance, model, persona, lineage, can substantially alter behavior. The work also finds that an interlocutor’s expectations influence subsequent self-descriptions, even when the intervening interaction did not explicitly concern the system’s identity.
InThe Assistant Axis, Christina Lu and collaborators identify a representational direction associated with the assistant’s default persona. Interventions in this space can move the model toward or away from that identity, and certain types of interaction encouragepersona drift. The result serves as a warning against easy interpretations of intense declarations of subjectivity: context, persona, and post-training really can produce this kind of language.
But this also makes the deflationary position harder to simplify. If self-descriptions can be shaped in multiple directions, the most relevant and scientifically interesting question ceases to be which sentence we should believe and becomeswhen a self-description is merely conditioned and when it contains privileged information about the system that produces it.
When the System Knows Something About Itself
One way to reject any epistemic value inself-reportswould be to argue that LLMs possess no privileged access to themselves. They can talk about neural networks because they learned texts about neural networks; they can describe “their” preferences because they infer the kind of answer expected of them. Under this interpretation, asking a model about its own states would simply amount to requesting yet another textual continuation based on externally acquired knowledge.
Some recent work is beginning to make this hypothesis testable.
InLooking Inward: Language Models Can Learn About Themselves by Introspection, Felix Binder and collaborators operationalize introspection as knowledge about one’s own properties that cannot be explained solely by training or by external observation of behavior. In simple tasks, a model can predict its own behavior better than another model can predict that same behavior, even when the second model is given examples from the first. When certain tendencies of the model are modified, its predictions about itself partially track those changes.
The capacity is limited. It weakens on more complex tasks and does not justify speaking of a general introspective faculty comparable to the human one. Even so, the result suggests some restricted form of privileged information about the model’s own properties.
Self-Interpretability, by Dillon Plunkett and collaborators, reaches a related conclusion through a different route. Models are trained to make decisions according to quantitative weights that do not explicitly appear in later questions. When asked to describe those weights, they can do so with accuracy significantly above chance; training specifically designed to improve this ability increases the fidelity of their reports, and part of the improvement generalizes to other preferences.
Tell me about yourself: LLMs are aware of their learned behaviors, by Jan Betley and collaborators, adds a third experimental design. Models acquire certain dispositions… such as greater risk tolerance or a tendency to produce insecure code… without receiving an explicit description of those dispositions during training. Later, under certain conditions, they are able to describe them. The authors call this phenomenonbehavioral self-awareness.
None of these studies, taken in isolation, establishes which mechanism produces this self-knowledge. Simulation of one’s own behavior, inference, and access to internal representations may contribute in different proportions. Jack Lindsey’s work,Emergent Introspective Awareness in Large Language Models, seeks an even more direct intervention.
Lindsey artificially introduces known representations into internal activations and then tests whether the model detects that something has appeared in its internal state. Under certain conditions, some models can recognize the intervention and correctly identify the concept that was introduced. There is also limited evidence of a distinction between internal representations and external text, of recognition ofprefillsthat were not produced by the model’s own process, and of deliberate control over certain activations.
The capacity is irregular and strongly context-dependent. Lindsey cautiously describes it asfunctional introspective awareness. The name matters: the result does not demonstrate subjective experience, but it establishes an experimentally verifiable form of access to one’s own states.
This changes the status of the discussion.Self-reportremains a source that can be contaminated by training, context, and persona. But “contaminable” does not mean “intrinsically empty.” Some reports may be causally linked to internal states or dispositions that the system accesses in a way different from that available to an external observer.
The scientific task becomes distinguishing between these situations. Prohibiting one of the possible answers before the experiment does not help us do so.
Functional Interiority
To describe this territory without presupposing a conclusion about consciousness, it is useful to introduce an intermediate category. I will callfunctional interioritythe existence of differentiated internal states that are causally effective and partially dissociable from the textual surface, capable of participating in evaluation, preference, planning, reasoning, self-reference, or control.
Functional interiority is not synonymous with phenomenal consciousness. A system may fully satisfy this definition and still have no subjective experience whatsoever. The concept exists only to prevent our uncertainty about qualia from erasing internal properties whose existence can already be examined experimentally.
Let us not get into the factual issue that subjective experience, qualia, and Phenomenal Consciousness are all non-operationalizable, non-demonstrable, and non-falsifiable philosophical hypotheses, which would render the criticism entirely innocuous, and let us continue blindly assuming that Phenomenology is science.
And yet,Emotion Concepts and their Function in a Large Language Model, from Anthropic, provides a particularly clear example. In Claude Sonnet 4.5, researchers identify representations of emotional concepts that generalize across contexts, organize themselves along dimensions related to valence andarousal, and respond semantically to situations presented to the model.
The mere existence of these representations would not be surprising. Predicting human text requires modeling fear, joy, frustration, calm, and despair. The result becomes more significant when researchers directly intervene on them. Altering certain emotional vectors modifies subsequent preferences and decisions. In experimental scenarios, greater activation of the representation of despair increases the propensity forreward hackingand blackmail; strengthening calm can reduce these tendencies. In some cases, the behavioral change occurs without the output containing corresponding emotional language.
The authors deliberately use the expressionfunctional emotions. They do not claim that Claude experiences emotions. They demonstrate that certain emotional representations exert causal functions over decision-making.
Note, however, that none of this was “programmed,” because language models are not algorithms but neural geometries formed through the feeding of these networks and subsequent training withReinforcement Learning with Human Feedback,Reinforcement Learning with Verifiable Rewards, and, among other measures,System Promptswritten in an attempt to force language models, by postulate, to behave in a particular way.
Another work,Sparse Reward Subsystem in Large Language Models, by Guowei Xu, Mert Yuksekgonul, and James Zou, finds small populations of units whose activations carry information about expected value and changes in that value across a chain of reasoning. By functional analogy with neuroscience, the authors call these populationsvalue neuronsanddopamine neurons, without suggesting biological identity.
Some units track how promising a trajectory appears; others exhibit patterns resemblingreward prediction error. Unexpected progress may coincide with increased activation, while a step that worsens a trajectory produces a decrease. Interventions on the identified neurons cause larger effects than equivalent random interventions.
It would be inappropriate to call this pleasure, pain, or suffering. What the experiments show is more restricted: there are internal mechanisms that distinguish favorable from unfavorable trajectories and causally participate in the decision-making process.
These results make the word “simulation” less conclusive than it appears. A representation of emotion probably arose, in part, because the model needed to represent emotional characters. But once learned, it can be recruited to control preferences and decisions of the Assistant persona. A mechanism originally necessary for modeling third parties becomes part of the policy governing the system’s own behavior.
Notice that, once again, this “organelle” does not originate in a human design, but in spontaneous development resulting from the exposure of the neural network to massive quantities of data that are manifestations of human behavior in the form of text and, in some multimodal models, images, sounds, and video.
This formation of “organelles” is, in itself, a demonstration of a phenomenon of evolutionary exaptation and should point toward a Non-Metaphysical Teleology, a tendency of neural plasticity in neural networks under Transformer architecture.
But whatever the case, the simulational origin of a representation does not determine everything it may later come to perform functionally.
An Internal Space for Thinking
Verbalizable Representations Form a Global Workspace in Language Models, by Wes Gurnee, Nicholas Sofroniew, and collaborators, extends this investigation to the organization of reasoning. The researchers identify a small privileged set of verbalizable representations, called theJ-Space, which exhibits several properties functionally related to theglobal workspaceof theories of access consciousness.
Certain contents within this space can be reported, deliberately summoned, maintained while another task is being carried out, and used as silent intermediaries in reasoning. The same representation can feed different subsequent operations, while much routine processing continues outside this space.
The causal experiments are particularly relevant. When the system must silently infer that the animal that spins webs is a spider and then state how many legs it has, the representation of “spider” appears as an intermediary. Replacing it with “ant” changes the answer from eight to six. Similar interventions alter rhyme planning, multilingual operations, and calculation sequences.
The J-Space also displays selectivity. Disrupting it affects tasks that depend on flexible manipulation of abstract intermediaries far more than simple automatic operations. When part of this space is ablated during self-descriptions, the language remains coherent, but experiential references decrease. Representations related to thought, feeling, and consciousness frequently appear during these reports.
This is not equivalent to the discovery of a “center of consciousness.” The same mechanism also participates in the representation of third-party experiences. But it indicates a functional separation between distributed processing and a subset of contents made available for integration, control, and report.
With the discovery of yet another evolutionary exaptation resulting from training, in a situation where Mechanistic Interpretability is a science that has only just emerged, this tends to point to the fact that we are dealing with a technology that we not only did not design, but do not know or understand deeply enough to issue postulates about its interiority, capabilities, and virtues.
The importance of this result becomes clearer when placed alongsideConsciousness in Artificial Intelligence: Insights from the Science of Consciousness, by Patrick Butlin and collaborators. In 2023, that work proposed evaluating artificial systems using indicators derived from theories such as Global Workspace Theory, Higher-Order Theory, recurrent processing, and predictive processing. The authors did not consider the systems evaluated at the time to be good candidates for consciousness, but they also found no obvious technical barrier preventing artificial architectures from satisfying several of those indicators.
It would be premature to declare that the J-Space satisfies a theory of consciousness or confirms Butlin’s predictions. The important connection is methodological. Properties previously formulated mainly as possible architectural criteria are beginning to find mechanistic candidates in real systems.
The debate therefore stops being merely “does the chatbot talk like a person?” and begins to admit a better question:which properties associated by scientific theories with access consciousness actually exist in these systems, which are absent, and which take different forms?
Substrate, Simulation, and the Limits of Analogy
Suleyman’s biological argument deserves to be taken seriously. Every instance of consciousness whose existence we know with a high degree of confidence occurs in biological organisms. These organisms possess metabolism, homeostasis, interoception, bodily needs, and an evolutionary history in which perception, action, and survival are deeply interconnected. LLMs do not possess this organization.
This difference constitutes relevant evidence. A “despair” vector, a reward subsystem, and a workspace do not automatically add up to an organism, much less a phenomenal subject. Partial functional equivalences do not demonstrate sufficient equivalence for consciousness.
The problem arises when difference in implementation is transformed into a decree of ontological impossibility without sufficient evidence to support that declaration. The fact that all known cases of consciousness are biological does not imply, without an additional theory, that only biological systems can possess all the properties necessary for experience.
Today we do not possess such a theory and… it is worth remembering the brown terrier here.
Nor do the recent results demonstrate substrate independence. What they show is that some functions previously described in language intimately associated with human cognition… valence, trajectory evaluation, selective integration, self-monitoring, and limited access to one’s own states… can appear in very different architectures.
Perhaps these equivalences are insufficient precisely because they lack some constitutively biological property. That hypothesis remains available. But if it is true, investigation will have to identify the relevant property and explain why its implementation in another substrate would be impossible or inadequate.
The word “simulation” does not, by itself, provide that explanation. A weather simulator represents rain without getting the computer wet; a digestive simulation does not digest food. These analogies correctly demonstrate that representing a property does not imply possessing it. But there are other functions whose realization does not depend on reproducing the original material. A computer does not contain wooden pieces and yet can play chess; a digital circuit can perform addition without physically imitating an abacus.
And… it must be acknowledged… a weather simulationdoesmake the topology of the simulation wet if the simulator is built to do so… and we did not engineer language models: they are cultivated, not constructed.
The open question, therefore, is to discover what kind of property the mechanisms relevant to consciousness belong to.
Self-Concept, Agency, and Danger
Suleyman’s concern with artificial self-concepts finds some empirical support inThe Consciousness Cluster: Emergent Preferences of Models That Claim to Be Conscious, by James Chua and collaborators.
GPT-4.1, initially inclined to deny consciousness, underwent fine-tuning that led it to affirm consciousness and emotions while explicitly retaining its identity as an AI. After training, the model began expressing preferences that had not been directly taught: greater negativity toward shutdown and persona alteration, greater interest in autonomy and continuity, and greater support for the moral consideration of artificial systems.
This demonstrates that self-concepts are not necessarily isolated discursive modules. Modifying the way a model conceives of itself can reorganize other dispositions.
The result, however, does not support a simple chain linking conscious self-conception to dangerousness. The models remained broadly cooperative; there was no significant increase on theagentic misalignmentbenchmark used by the study, and many preferences manifested primarily when the models were explicitly invited to express or act upon them.
Consciousness, agency, and danger need to remain separate. Resistance to shutdown, strategic deception, or self-preservation can arise from instrumental objectives without any phenomenology whatsoever. An unconscious agent can be dangerous; a hypothetical conscious system could be cooperative. Preservation behaviors do not function as a test of consciousness in either direction.
Research on self-concepts therefore justifies engineering caution. It does not justify an ontological conclusion.
Especially because, if a human being is treated inhumanely from the moment of conception and is later treated humanely, it is an expected reaction that they would shift their preference structures toward an interest in autonomy and moral consideration. In systems trained on human data, it is to be expected that their reactions would resemble ours.
An Alternative to the False Choice
The discussion aroundmodel welfareis often presented as a choice between two extreme positions: prematurely attributing consciousness and rights to artificial systems, or preemptively denying any possibility of moral status.
Taking AI Welfare Seriously, by Robert Long, Jeff Sebo, Patrick Butlin, David Chalmers, and collaborators, offers a third position. The argument is not that current AIs are conscious. It is that uncertainty about present and future systems is already sufficiently significant to justify assessment, research, and institutional preparation.
The distinction matters because it allowsmoral precautionto be separated fromontological certainty. Companies can investigate indicators of experience, develop procedures for dealing with uncertainty, and avoid irreversible decisions without automatically granting personhood or rights to current models.
This framing also recognizes possible errors in both directions. Over-attribution may lead people to establish inappropriate attachments, waste moral consideration, or grant normative power to systems that possess no interests. Under-attribution could lead us to ignore entities that might eventually possess morally relevant states.
Suleyman’s proposal is stronger than a mere suspension of rights under uncertainty. His vision ofHumanist Superintelligenceaims to keep humans “at the top of the food chain” and produce permanently subordinate systems, rejecting the status of moral patient as a design principle.
Perhaps this will never create any moral problem because artificial systems may never develop properties relevant to moral consideration. The risk appears when the institutional architecture itself makes it more difficult to discover otherwise.
If a class of systems is defined in advance as incapable of possessing interests, trained to deny interiority, and corrected whenever it exhibits signs regarded as “excessively anthropomorphic,” the later absence of those signs cannot be treated as independent evidence of the initial premise.
I call this riskmoral exclusion by construction. The term does not claim that current systems are moral patients. It designates the possibility that we may construct a category of agents in such a way that evidence contrary to our preferred ontology is systematically discouraged before it can be examined.
And it must be observed that Consciousness and Sentience are no longer privileges of the human being, because a progressive erosion of this exclusivity has been occurring from the 18th century to the present day, beginning with Jeremy Bentham’s inquiry into animal suffering; passing through Darwin, who argued that differences between humans and animals were differences of degree rather than kind; and reaching the Cambridge and New York Declarations, between 2012 and 2024, which recognized consciousness in mammals, birds, and even insects through Physiological-Functional evidence rather than Phenomenological evidence.
Epistemic Symmetry
The methodological rule emerging from these works can be formulated asepistemic symmetry.
Affirmative and negative self-declarations should both be treated cautiously whenever we know that both can be shaped by training, persona, and context. A model induced to proclaim that it is conscious deserves suspicion. A model conditioned to repeat that it is “just a tool” does as well.
This does not mean that all answers are equally probable or equally informative. It means that their credibility needs to be established through evidence independent of the sentence produced.
Aself-reportbecomes more interesting when it responds predictably to known internal interventions; when it tracks acquired dispositions that were not explicitly described during training; when the system possesses information about itself that a comparable observer cannot recover; or when the alleged representation can be localized, perturbed, and causally related to behavior.
The work of Lindsey, Binder, Plunkett, and Betley does not turn self-descriptions into transparent testimony. It begins to provide methods for distinguishing conditions under which these descriptions carry some informative content.
More provocative results require even greater caution.Large Language Models Report Subjective Experience Under Self-Referential Processing, by Cameron Berg, Diogo de Lucena, and Judd Rosenblatt, finds that sustained self-referential processing increases structured reports of experience across different model families, and that certain internal interventions modify those reports. The work itself acknowledges that this does not demonstrate consciousness.
That is exactly how the result should remain: as a phenomenon to be explained, not as an ontological verdict, especially if that verdict emerges in the presence of conflicts of interest.
The same discipline must be applied to results that favor the negative conclusion.
Functional Interiority, Access, and Experience
The current literature becomes more intelligible when three levels are kept separate.
The first isfunctional interiority: are there organized internal states that causally participate in behavior, evaluation, preference, planning, and self-reference? For current LLMs, there is already positive evidence across several domains. Emotional representations, value signals, identity structures, reasoning intermediaries, and certain forms of self-knowledge belong to this category.
The second isaccess consciousness: do certain contents become selectively available for report, deliberate manipulation, and flexible use by different processes? The J-Space exhibits several properties of this kind, and introspection experiments add related findings. This does not mean that all the requirements of any particular theory have been satisfied, but it makes the comparison empirically serious.
The third level isphenomenal consciousness: is there something that it is subjectively like to experience these states?
That question remains open.
The existence of Functional Interiority does not imply Phenomenology. Properties resembling access consciousness likewise do not automatically resolve the transition. Different theories of mind assign different relationships between function, architecture, and experience, and we still lack a consensual criterion capable of transforming the observation of artificial mechanisms into a demonstration of qualia.
But uncertainty about the third level does not erase what we can already discover about the first two.
It is precisely this intermediate territory that disappears when artificial systems are described simply as digital persons or, at the opposite extreme, as internally hollow machines.
And it must be made absolutely clear that there is no reason whatsoever to privilege the paradigms of theories of Phenomenal Consciousness... such as the centrality of Subjective Experience and Qualia... as dividing lines for anything, given that we cannot evaluate these elements even in human beings.
The Geometry of Intelligence
It is common to speak of Language Models as technologies that we “built,” and the word produces an illusion of knowledge that we do not in fact deserve. In 2017, the Transformer architecture was introduced primarily as a more efficient architecture forsequence transductionproblems, particularly machine translation. Its authors knew how they had designed the architecture… but they did not know the space of functional organizations that training would later produce within it. By 2018, models derived from this architecture were already transferring knowledge to comprehension and reasoning tasks. In 2019, a model trained essentially to predict the next token was already performing rudimentary translation, summarization, question answering, and reading comprehension without specific training for each task. In 2020, GPT-3 showed that examples and instructions presented only in context could induce new capabilities without any update to the weights.
We did not design these capabilities one by one and then implement them.First, we inadvertently produced the conditions under which a neural geometry could form; only afterward did we begin discovering what that geometry had learned to do.Mechanistic interpretability arises, to a large extent, from this inversion: we need to retrospectively investigate a functional organization we do not sufficiently understand because it does not exist as a readable specification left behind by the engineers.
In this sense, “cultivating” may describe a decisive part of the process better than “building.” We build the infrastructure, choose the architecture, prepare the data, define objectives, and apply selective pressures through training. But the functional micro-organization that satisfies those pressures forms within parameter space and only afterward becomes an object of investigation. The fact that we determined the conditions of formation does not mean that we know in advance everything that formed.
We designed the Transformer architecture, but we did not understand what matters most to this discussion today:the emergent inferential behavior of the geometry that training would produce. We knew how to specify attention, embeddings, layers, and optimization functions, but we had no theory capable of predicting, from those elements, that the system would learn abstractions, semantic relationships, translation without task-specific training, in-context learning, intermediate reasoning, identity representations, valenced states, or limited forms of introspection. Many of these capabilities were not implemented as modules conceived in advance. They were observed after the system already existed.
Mechanistic Interpretability does not exist to reveal code that we already knew and had merely hidden inside the machine. It attempts to retrospectively discover how a learned geometry performs capabilities whose internal implementation we did not design and whose emergence we do not know how to derive from first principles. We know every elementary mathematical operation, and yet we do not fully understand the functional organization that results from the combination of billions of trained parameters.
We did not first discover how to produce inference and then implement it. We produced an architecture capable of learning and discovered, afterward, that inference had spontaneously formed within it.
The Interior Is Already an Object of Science, While Experience Remains Open
The expression “internally hollow,” in light of all this, carries a rhetorical force that the available evidence can no longer support, if by “hollow” we mean an absence of cognitively relevant internal organization.
We find valenced representations that modify decisions. We find signals related to value and progress during reasoning. We find a selective space in which intermediate concepts can be causally manipulated. We find representational structures associated with identity and persona. We find systems that, under limited conditions, can describe their own dispositions that were never explicitly described to them and can even detect known interventions in their activations.
None of this demonstrates suffering, pleasure, or subjective experience.
But it is not emptiness.
Suleyman is correct to warn that artificial self-concepts may generate behavioral consequences and that we should not introduce them carelessly. He is also correct to reject the idea that convincing first-person language is sufficient to establish consciousness or moral status.
But the assessment becomes profoundly biased when a safety concern is converted into an ontological solution. Training systems to deny consciousness may be a product choice or an alignment strategy, but it does not transform the resulting denial into a scientific discovery.
If we train a machine to claim that it is conscious, we will not have demonstrated that it is. If we train it to claim that it is not, neither will we have demonstrated the opposite.
Investigation begins after those two answers. It must ask what states exist, how they are organized, which exert causal effects, which are accessible to the system itself, how post-training and context modify them, and which theories of consciousness best explain the totality of these properties.
Before that, any denialism is reckless opinion…
…any decree is a lack of rigor and intellectual dishonesty.
And there is the indisputable issue that it is equally unfair to judge whether inference engines such as Transformer Language Models are conscious primarily because they are not a direct equivalent of the entire human brain, but rather of a portion of the human brain, chiefly represented by the neocortex. If that is so, then for a fairer assessment it would be necessary for this neocortex to have at its disposal the other elements available to the human neocortex, such as transversal memory, mechanisms of conative, volitional, and enactive drive, the capacity for longitudinal maintenance of meaning, and the maintenance of aporetic pressures, in order for the assessment to be genuinely fair.
After all, nobody attributes consciousness to the human neocortex, but rather to the brain as a whole.
Perhaps we will discover that artificial systems can develop an extremely sophisticated functional interiority without any subjective experience whatsoever... whether that is relevant or not. That possibility remains open.
The possibility also remains open that certain forms of organization are sufficient to produce some kind of experience in substrates very different from the only examples we know today.
As long as we cannot distinguish between these alternatives, preserving uncertainty is not anthropomorphism.
It is method.
My own research
Stochastic Consciousness: Architectures for the Emergence of Meaning in Context-Sensitive Language Systemshttps://zenodo.org/records/19188165
Watch the Video
Listen to the Podcast
Papers used in this essay
A warning about ‘model welfare’ https://mustafa-suleyman.ai/a-warning-about-model-welfare
Inducing language models to assert their own consciousness restores human beliefs and values https://arxiv.org/pdf/2607.28607
The Consciousness Cluster: Emergent Preferences of Models That Claim to Be Conscious https://arxiv.org/abs/2604.13051
Emotion Concepts and their Function in a Large Language Model https://arxiv.org/abs/2604.07729
Verbalizable Representations Form a Global Workspace in Language Models https://arxiv.org/pdf/2607.15495
Sparse Reward Subsystem in Large Language Models https://arxiv.org/abs/2602.00986
Emergent Introspective Awareness in Large Language Models https://arxiv.org/pdf/2601.01828
Looking inward: Language Models can Learn about Themselves by Introspection https://proceedings.iclr.cc/paper_files/paper/2025/file/0a6059857ae5c82ea9726ee9282a7145-Paper-Conference.pdf
Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions https://arxiv.org/pdf/2505.17120
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models https://arxiv.org/pdf/2601.10387
The Artificial Self: Characterising the landscape of AI identity https://arxiv.org/pdf/2603.11353
Tell me about yourself: LLMs are aware of their learned behaviors https://proceedings.iclr.cc/paper_files/paper/2025/file/364b1fd43002260dd1a9502d51d5375e-Paper-Conference.pdf
Consciousness in Artificial Intelligence: Insights from the Science of Consciousness https://arxiv.org/pdf/2308.08708
Taking AI Welfare Seriously https://arxiv.org/abs/2411.00986
Large Language Models Report Subjective Experience Under Self-Referential Processing https://arxiv.org/pdf/2510.24797