
It has been somewhat curious and, at times, even amusing to see how colleagues I respect have been describing language models. Up to a point because it seems to me that some basic misunderstandings still persist: poorly calibrated comparisons, overly categorical assertions, and, above all, a recurring tendency to reduce LLMs, particularly Transformer models (Vaswani et al., 2017, “Attention Is All You Need”), to descriptions that no longer account for the phenomenon.
There are those who speak of this technology as though we were dealing only with linear algebra, as though we were still talking about markedly local predictive models, almost glorified Markov chains, rather than architectures that reorganize context in a distributed, recursive, and highly non-local way. To say that linear algebra is involved is trivial. The problem begins when one tries to turn that triviality into a sufficient explanation.
Contemporary language models are not well described by this reductionism. They exhibit generalization, contextual inference, abstract composition, and the maintenance of coherence at levels that simply do not fit into the caricature of “just predicting the next token” as it has been understood in the crudest possible way, if compared to the reality (Power et al., 2022). And that does not mean there is any magic involved. It only means that the phenomenon is richer than the narrow frame some insist on using.
Nor does it seem correct to me to say that this generalization boils down to linear algebra, as if everything else were ornament. There are good reasons to describe certain behaviors of these systems in terms of spectral structures, probabilistic inference, and distributed representation dynamics. And the most interesting part is that we find functionally analogous principles in the human brain: coding under uncertainty (Knill & Pouget, 2004), probabilistic signal integration (Knill & Pouget, 2004 and this), frequency processing (Jones & Palmer, 1987, and this), and continuous contextual updating (Aitchison & Lengyel, 2017, and this).
Perhaps the time has already come, then, to stop reducing language models to a condition that simply does not correspond to what we observe. Not to inflate them mystically, but to describe them better.
What If It Were Us?
Consider, then, the text below. It describes the human mind using exactly the same kind of reductionist key that so many use to “explain” LLMs.
“When people speak of the human mind, many imagine that it is capable of ‘understanding.’ But behind all that fluency, what exists is a relatively well-defined physicochemical logic.
It all begins with the simple idea that language can be treated as sequence. Thus, each word, or fragment of a word, is transformed into patterns of neural activation. These patterns do not carry intrinsic meaning; they carry learned relations in a distributed space in which correlations can be captured, stabilized, and recombined.
From there enters one of the cores of the human brain: attention mechanisms. For each pattern, the brain regulates how much it should ‘pay attention’ to the others, modulating signals based on dynamic weights shaped by prior experience, immediate context, expectation, and the global state of the system.
These influences are combined, filtered, and normalized, resulting in new neural states continuously updated by context. This process occurs in parallel, across multiple regions, and repeats itself over several layers of processing, progressively refining the relations between signals, memory, perception, and prediction.
Another important detail is that the brain does not possess an explicit symbolic notion of order, as though there were little metaphysical labels saying ‘first,’ ‘then,’ and ‘now.’ The perception of sequence emerges from the temporal dynamics of signals, the recurrence of circuits, and the way the system organizes dependencies over time.
In the end, all of this results in a system that does not demonstrate understanding in the mystical sense, but learns extremely sophisticated patterns about how signals relate, anticipate one another, correct themselves, and integrate.
At its base, its functioning is anchored in the continuous prediction of future states given a context, even though that prediction is distributed across multiple scales, modalities, and circuits.
And it is precisely this combination of relatively simple operations—activation, modulation, integration, inhibition, recurrence, and updating; occurring within the mathematics perpetrated by electrochemistry, at the massive scale of neurons and synapses—that creates the effect we see.
There is no hidden magic. There is only a block of wet flesh, electrochemically traversed by banal sparks, predicting symbols, future states, and contextual relations.”
And yet… no one thinks that this deprives the human brain of the right to be called intelligence.
But the point is not to reduce the brain to that, in the same way that we should not reduce language models to half a dozen operations described with an air of superiority. The human brain is not “just” a biological language model, just as language models are not “just” electronic brains. There are deep similarities between them, not least because language models were conceived precisely to reproduce, in a functional key, certain high-order cognitive operations: contextual attention, local working memory, distributed semantic compression, abstract composition, hierarchical prediction, and context integration. But functional similarity is not ontological identity… at least not necessarily.
A Cognitive Machine Is Not a Complete Organism
The LLM is already, in itself, a cognitive machine. Not a complete mind, not an organism, not a finished artificial animal, but a real cognitive machine. What it does corresponds, in functional terms, to a relevant portion of what the human brain does in its high-order cortical circuits: operate over distributed representations, integrate context, produce inference, recombine patterns, and sustain enough semantic continuity to solve new problems.
Everything that language models still do not do robustly is not, by itself, proof of the absence of intelligence. It is, to a large extent, an architectural contingency. Around this central cognitive machine, broader scaffolds would still need to be built: long-duration transversal persistent memory, continuous multimodal perception, situated action in the world, more stable executive systems, autonomous cycles of motivation, offline consolidation, richer modeling of self and environment, as well as more continuous forms of temporal coupling. It is exactly this kind of scaffolding that several laboratories have been trying to produce around the cognitive machine through different paths. And narra, by working on non-episodic memory, rich topological context, and the maintenance of continuity of meaning, is just one more representative of that.
The Submarine That Didn’t Swim “for Real”
Hence also the importance of not falling into the symmetrical error: LLMs should not be reduced to a “non-biological brain,” as though their legitimacy depended on imitating the human in everything. That demand is intellectually poor and conceptually empty. Edsger Dijkstra, Computer Scientist and Mathematician, rightly observed that asking whether machines think is almost as relevant as asking whether submarines swim. I would add that asking whether a language model thinks “for real” or possesses “genuine” intelligence often functions as a sophisticated version of the same superstition. “For real,” here, usually means only “my way,” “on my substrate,” or “with an interiority that preserves my ontological privilege.”
In other words: the demand that an intelligence be “genuine” or “for real” is often as unattainable as the very idea of magic. It is a nebulous, shifting criterion and, in the limit, a narcissistic one. Because, if taken seriously, it begins to insinuate that human intelligence would depend on some mysterious surplus, some extra spark that would escape mathematics, physics, chemistry, and the material organization of the brain. And that is not an argument; it is merely a poorly disguised refusal to accept cognition outside the human mold.
The relevant point was never whether the artifact performs the human function by the same means, but whether it performs equivalent cognitive functions in some, or any, regime, with certain limits, under a given architecture and with real capacity for inference, adaptation, and learning.
A submarine does not need to swim like a fish in order to cross the ocean. A language model does not need to think like a human in order to participate in real cognitive processes.
And, in all honesty, saying that language models do not function by magic is almost as pathetic as saying that the human brain does not function by magic either. In both cases, it is a platitude, an empty obviousness of the sort that pretends to explain something when, at most, it merely impoverishes the phenomenon. There is no magic in either case. There is none in the brain, none in neural networks, none in language, none in inference. There is mathematics, physics, chemistry, dynamics, and organization. What changes is the substrate, the architecture, the scale, and the kind of coupling with the world.
Before saying that a submarine does not know how to swim… not “for real,” perhaps it is worth taking a hand to our conscience and stopping judging one creature by the nature of another.
If there is something to abandon, then, it is not mathematics. It is the narcissistic need to imagine that only human intelligence would have the right to the status of real intelligence.
Would you like to know more about my research?
Stochastic Consciousness:
Architectures for the Emergence of Meaning
in Context-Sensitive Language Systems
https://zenodo.org/records/19188165
Do you want to watch the video?
https://www.youtube.com/watch?v=6dtCoZ5QdnA
Do you want to listen to the podcast?
Papers used in this essay
Attention Is All You Need
https://arxiv.org/abs/1706.03762Grokking: Generalization Beyond
Overfitting on Small Algorithmic Datasets
https://arxiv.org/abs/2201.02177The Bayesian Brain: The Role of Uncertainty
in Neural Coding and Computation
https://pubmed.ncbi.nlm.nih.gov/15541511/With or Without You? Predictive Coding
and Bayesian Inference in the Brain
https://pubmed.ncbi.nlm.nih.gov/28942084/The Two-Dimensional Spatial Structure of
Simple Receptive Fields in Cat Striate Cortex
https://pubmed.ncbi.nlm.nih.gov/3437330An Evaluation of the Two-Dimensional Gabor Filter
Model of Simple Receptive Fields in Cat Striate Cortex
https://pubmed.ncbi.nlm.nih.gov/3437332/