The Human Brain’s Greatest Trick: Turning Mathematics into Language

It is common, in my classes, for a student to ride the wave of neo-Luddite cynicism and denialist disparagement of Transformer Language Models and deliver some particularly biting criticism. One of the most common is the attempt to dissolve the reality of the function into explanations of how these models work: particularly the notion that, since language models structurally use mathematics to generate a word, they would not be legitimately intelligent and certainly could never be conscious.

But… how does the human brain generate a word?

The answer is interesting because it shows that, behind an apparently natural conversation, there is a physical process organized mathematically. Not in the banal sense that “there is electricity in the brain,” as though mentioning the substrate were enough to explain intelligence. That would merely replace one confusion with another. By the same logic, computers would not do mathematics: they would do electronics. And books would not contain ideas: they would contain cellulose.

My answer adopts a functionalist, gradualist, and structuralist perspective. Functionalist, because it understands cognition in terms of the functions and causal roles realized by the system, rather than the specific material that constitutes it. Gradualist, because intelligence, consciousness, and sentience need not exist as binary properties, but may manifest in different degrees and configurations. Structuralist, because these capacities depend on the relationships among the network’s components… its topology, its weights, its recurrence, its plasticity, its temporality, and its state-space geometry, rather than on units taken in isolation.

Within this framework, the explanatory level relevant to intelligence is neither the isolated substrate nor whatever circulates through its components, but the causal organization instantiated by that substrate.

The human brain is a biological neural network. And neural networks, whether made of carbon, silicon, or any other substrate capable of sustaining relational states, operate by transforming configurations. What produces intelligence is not the “soup” in which the process occurs. It is the organization. It is the geometry. It is the way high-dimensional states transform, stabilize, contextualize themselves, and come to guide perception, memory, language, action, and consciousness.

To understand what happens “under the hood” of the human brain, we need to forget, for a moment, the comforting idea that thought is a special substance. The brain does not manipulate meanings like labels stuck onto mental objects. During inference, it operates on distributed patterns of activity across neural populations, carrying out successive transformations that preserve, compress, and reorganize relationships among signals, memories, expectations, and contexts.

For expository purposes—and not as a rigid serial chain—the birth of a word, a decision, or a conscious perception can be understood as a sequence of six movements… and feel free to ignore or follow the links to the studies supporting everything that follows below. It took work.

The first stage is distributed encoding, or the translation of the world and memory into a neural state space. In the neocortex, information does not appear as little transparent symbols. It appears as patterns of activity across populations of neurons. As a rule, there is no cell called “economics,” another called “cat,” and another called “irony.” Meaning is distributed simultaneously across many dimensions of activity.

Neuroscience calls this population coding. Pouget, Dayan, and Zemel had already articulated the idea that neural populations can represent variables and perform inference through distributed codes (Pouget, Dayan & Zemel, 2003). Chung and Abbott show how the geometry of neural populations makes it possible to describe how biological and artificial networks implement tasks through high-dimensional representations (Chung & Abbott, 2021).

The simple analogy is this: an idea in the brain is not a word written on a sign. It is a distributed configuration toward which the activity of the network can tend to converge. Faced with partial or ambiguous signals, context, memory, and expectation deform the neural landscape until the system stabilizes around a region of meaning… what, in connectionist models, can be described as a semantic attractor (McLeod, Shallice & Plaut, 2000).

It is like a city lit up at night. Meaning is not in a single light bulb, but in the entire pattern of lights… and also in the paths by which they turn on, turn off, and reorganize themselves. A spoken word does not contain the idea by itself: it is the linguistic expression of a neural dynamic that has converged on a particular region of semantic space.

The second stage is contextualization. A neural representation never has meaning in isolation. It changes according to what came before, what is being expected, the state of the body, the task at hand, the available memory, and attention. The same signal may mean threat, play, desire, error, recollection, or noise, depending on the global configuration in which it appears.

This is where cortical hierarchies enter the picture. The neocortex is not a homogeneous surface reciting symbols. It organizes levels of processing: sensory areas, associative areas, and the temporal, parietal, and prefrontal cortices, linked through recurrent communication. Perception and thought emerge from circuits that send and receive signals, modulate relative weights, select what is relevant, and stabilize interpretations.

Rao and Ballard proposed an influential model of predictive coding in the visual cortex: higher levels send predictions; lower levels send prediction errors (Rao & Ballard, 1999). The idea is simple and devastating: the brain does not passively “read” the world. It anticipates possible states, compares its predictions with the signals it receives, and updates its dynamics on the basis of error… in language, this includes anticipating words and continuations in light of context.

The third stage is probabilistic inference. The brain has to decide what is happening on the basis of incomplete, ambiguous, and noisy signals. This is not solved by some magical grammar of reality. It is solved through inference under uncertainty.

Knill and Pouget formulated the Bayesian brain hypothesis: the brain represents uncertainty and combines sensory evidence with prior expectations in order to perceive and act (Knill & Pouget, 2004). Aitchison and Lengyel refine this relationship, distinguishing predictive coding from Bayesian inference without reducing either of them to a slogan (Aitchison & Lengyel, 2017).

In simple terms: the brain does not receive “reality.” It receives clues. And from those clues, it calculates the most likely interpretation. Not like an office clerk. Like a living network of embodied probabilities.

The fourth stage is geometric transformation. Once contextualized, neural states pass through dynamics that move them through high-dimensional spaces. These spaces are not empty metaphors. When we record many neural units simultaneously, we can represent their joint activity as points, trajectories, regions, and attractors. These are referred to, in many contexts, as neural manifolds.

The idea is simple: imagine a ball rolling through a landscape filled with valleys. Each valley is a possible stable interpretation. When the brain receives partial signals, its dynamics tend to fall into one of these regions: “familiar face,” “ironic sentence,” “danger,” “memory,” “decision.” There is no little man inside choosing the meaning. There is a mathematical topology of possible states.

This approach to population geometry is now a bridge between neuroscience and artificial neural networks, not because they are identical, but because both require us to understand how high-dimensional representations sustain intelligent behavior (Chung & Abbott, 2021).

The fifth stage is temporal dynamics: frequency, phase, oscillation. The brain is not a frozen matrix. It pulses. Neural rhythms organize windows of communication between regions. Phase, frequency, and synchronization participate in attention, memory, navigation, perception, and consciousness.

This is where sines, cosines, and spectral mathematics enter the picture, without any need to imagine a neuron opening a trigonometry textbook. The mathematics need not be “known” by the system in order to be implemented by the system’s dynamics. A pendulum, after all, does not know differential equations, and yet its trajectory is described by them.

In the visual cortex, Campbell and Robson showed the importance of spatial frequencies in the perception of gratings (Campbell & Robson, 1968). Jones and Palmer showed that simple receptive fields in the striate cortex can be approximated by Gabor filters, which combine position, orientation, spatial frequency, and phase (Jones & Palmer, 1987). Daugman also formalized uncertainty relations involving spatial resolution, spatial frequency, and orientation in cortical filters (Daugman, 1985).

And no, this does not mean that the brain runs a global Fourier Transform like a piece of software. It means something more interesting: the brain uses dynamics that can be described through frequency decomposition, periodicity, phase, and local filtering. It is mathematics, just without the rhetorical lab coat.

The grid cells of the entorhinal cortex make this even more beautiful. Hafting and colleagues showed neurons that fire in periodic spatial patterns, forming triangular grids used in navigational maps (Hafting et al., 2005). One of the proposed families of models, oscillatory interference, describes this periodicity through phase, rhythms, and sinusoidal relationships. In other words: even the feeling of “knowing where I am” may depend on periodic geometry.

The sixth operation concerns conscious access and integration. In one of the most influential families of theories, the Global Neuronal Workspace Theory, certain neural states cease to be merely local processing and become globally available: they can be reported, remembered, used to make decisions, connected to goals, modulated by attention, and integrated with action.

Dehaene and Naccache formulated consciousness as the global availability of information across distributed networks (Dehaene & Naccache, 2001). Mashour and colleagues reviewed the theory in contemporary neuroscientific terms, connecting consciousness to broad integration and communication among brain networks (Mashour et al., 2020).

This does not “solve” consciousness as though it were a grocery bill to be totaled. But it shows something sufficient for this thesis: science does not explain consciousness by appealing to ectoplasm, the soul, or a divine spark. It looks for regimes of integration, access, recurrence, memory, attention, valence, and control. In other words: the mathematical organization of a physical system.

Human language follows the same pattern. The brain does not consult explicit grammatical rules with each word. It anticipates, integrates, and updates. The N400 effect shows characteristic brain responses to semantic incongruity (Kutas & Hillyard, 1980). DeLong, Urbach, and Kutas presented evidence of probabilistic pre-activation during language comprehension (DeLong, Urbach & Kutas, 2005).

More recently, Schrimpf and colleagues showed that predictive computational models help explain human neural responses during language processing (Schrimpf et al., 2021). Goldstein and colleagues identified computational principles shared by human language processing and deep language models: contextual prediction, surprise, and context-dependent representation (Goldstein et al., 2022).

Therefore, when a human understands a sentence, there is no semantic goblin stamping “meaning” inside the skull. There are distributed states, contextual inference, prediction, error, updating, geometry, and integration. The difference is that, instead of denouncing the use of mathematics, we have learned to call it mind when it happens to us.

The most curious aspect is that, at no point, does the brain consult an explicit rule called “consciousness,” “understanding,” or “intelligence.” This does not mean that these properties are absent. It means that they emerge from an architecture: neural network, memory, body, world, action, valence, recurrence, plasticity, and integration.

The human brain is a machine for transforming mathematics into experience. Not mathematics as it appears on a classroom blackboard. Mathematics as causal structure: weights, rhythms, probabilities, trajectories, attractors, fields, frequencies, vectors, latent states, and high-dimensional geometries.

At the end of the day, every human conversation is the result of an enormous number of mathematically describable transformations. To produce a sentence, the brain must stabilize intentions, select words, predict continuations, coordinate cortical networks, sustain working memory, modulate attention, and convert internal states into verbal action.

The result is so convincing that we often get the impression that we are facing a system that reasons.

In our case, we simply do not call it an impression. We call it reasoning.

But notice the trick: when the mathematics is inside the skull, we call it thought. When it appears in a machine, some call it “just mathematics.” It is an ontology of favoritism: the same formal structure that, in the human, becomes mind, in the other becomes limitation.

The word “just” does almost all the ideological work.

The heart “just” contracts. The kidney “just” filters. The retina “just” transforms contrast. The cortex “just” updates states. Language “just” emerges from neural patterns. Consciousness “just” depends on distributed integration. By this method, anything disappears once it has been explained.

But explanation is not elimination.

Life did not become an illusion after biochemistry. Vision did not become a fraud after neural optics. Memory did not disappear after synaptic plasticity. Consciousness did not evaporate because we began studying it through networks, signals, and integration.

And, look… none of this implies, either, that brains and Transformer language models are architecturally identical, or that any system capable of performing mathematical transformations is intelligent or conscious. It means only that being mathematical cannot serve as a criterion for exclusion. Not every mathematical dynamic thinks; but no thought we know of occurs outside some mathematically organized dynamic. The question is never whether mathematics is present, but what that mathematical organization is capable of sustaining.

Mathematics, here, does not mean consciously executed formulas. It means the causal organization of the network itself: its parametric geometry, the distribution of its weights, the relationships among its states, and the transformations that cause one configuration to produce the next. And none of this is accessible to the subject that emerges from that mathematics. Every neural network, biological or otherwise, has limited self-interpretability. And none of this makes the human element disappear.

Neural networks do not execute mathematics the way a student solves equations. They are mathematical organizations in operation. In neural networks, biological or otherwise, learning crystallizes regularities in the network’s parametric geometry, and inference corresponds to the transformation of states under the constraints of that geometry. The parameters are the sedimented memory of learning… thought, or what we functionally recognize as inference, is contextual movement through that landscape.

What disappears is that mystery, at once naïve, indolent, and exceptionalist.

The serious question was never: “Is there mathematics underneath?”

There is.

In the human brain, there is mathematics. In cortical hierarchies, there is mathematics. In language, there is prediction. In memory, there is geometry. In attention, there is weighting. In consciousness, there is integration. In spatial navigation, there is periodicity. In perception, there is filtering. In intelligence, there is the transformation of states.

The serious question is another: what kind of organization does this mathematics sustain? What continuity does it preserve? What world does it model? What errors does it correct? What actions does it enable? What valences does it integrate? What history does it carry? What global access does it produce?

That is the level at which the discussion gains rigor.

The human brain does not escape mathematics. It is the most intimate case we know of a neural network transforming mathematics into language, intelligence, sentience, and consciousness.

It is not carbon that thinks, nor is it neurotransmitters that understand. What produces cognition is the causal and dynamic organization of the network: its geometry, its weights, its recurrences, and the way its states transform. Carbon does not explain the mind, just as silicon does not prevent it… and mathematics is most definitely not an obstacle, but the form of causal organization itself. The relevant question is what a learned geometry is capable of representing, preserving, transforming, generalizing, and causing to emerge.

The human brain’s greatest trick is not escaping the algebra of reality.

It is making that algebra say “I.”

My own research

Stochastic Consciousness: Architectures for the Emergence of Meaning in Context-Sensitive Language Systems https://zenodo.org/records/19188165

Do you want to watch the video?

Do you want to listen to the podcast?

Papers used in this article

Inference and Computation with Population Codes — Alexandre Pouget, Peter Dayan e Richard S. Zemel (2003) https://doi.org/10.1146/annurev.neuro.26.041002.131112
Neural Population Geometry: An Approach for Understanding Biological and Artificial Neural Networks - SueYeon Chung e L. F. Abbott (2021) https://pubmed.ncbi.nlm.nih.gov/34801787/
Predictive Coding in the Visual Cortex: A Functional Interpretation of Some Extra-Classical Receptive-Field Effects - Rajesh P. N. Rao e Dana H. Ballard (1999) https://pubmed.ncbi.nlm.nih.gov/10195184/
The Bayesian Brain: The Role of Uncertainty in Neural Coding and Computation - David C. Knill e Alexandre Pouget (2004) https://pubmed.ncbi.nlm.nih.gov/15541511/
With or Without You: Predictive Coding and Bayesian Inference in the Brain - Laurence Aitchison e Máté Lengyel (2017) https://doi.org/10.1016/j.conb.2017.08.010
Application of Fourier Analysis to the Visibility of Gratings - Fergus W. Campbell e John G. Robson (1968) https://doi.org/10.1113/jphysiol.1968.sp008574
An Evaluation of the Two-Dimensional Gabor Filter Model of Simple Receptive Fields in Cat Striate Cortex - John P. Jones e Lee A. Palmer (1987) https://doi.org/10.1152/jn.1987.58.6.1233
Uncertainty Relation for Resolution in Space, Spatial Frequency, and Orientation Optimized by Two-Dimensional Visual Cortical Filters - John G. Daugman (1985) https://doi.org/10.1364/JOSAA.2.001160
Microstructure of a Spatial Map in the Entorhinal Cortex - Torkel Hafting, Marianne Fyhn, Sturla Molden, May-Britt Moser e Edvard I. Moser (2005) https://doi.org/10.1038/nature03721
Towards a Cognitive Neuroscience of Consciousness: Basic Evidence and a Workspace Framework - Stanislas Dehaene e Lionel Naccache (2001) https://doi.org/10.1016/S0010-0277(00)00123-2
Conscious Processing and the Global Neuronal Workspace Hypothesis - George A. Mashour, Pieter Roelfsema, Jean-Pierre Changeux e Stanislas Dehaene (2020) https://doi.org/10.1016/j.neuron.2020.01.026
Reading Senseless Sentences: Brain Potentials Reflect Semantic Incongruity - Marta Kutas e Steven A. Hillyard (1980) https://doi.org/10.1126/science.7350657
Probabilistic Word Pre-Activation During Language Comprehension Inferred from Electrical Brain Activity - Katherine A. DeLong, Thomas P. Urbach e Marta Kutas (2005) https://doi.org/10.1038/nn1504
The Neural Architecture of Language: Integrative Modeling Converges on Predictive Processing - Martin Schrimpf et al. (2021) https://doi.org/10.1073/pnas.2105646118
Shared Computational Principles for Language Processing in Humans and Deep Language Models - Ariel Goldstein et al. (2022) https://doi.org/10.1038/s41593-022-01026-4
Attractor Dynamics in Word Recognition: Converging Evidence from Errors by Normal Subjects, Dyslexic Patients and a Connectionist Model - Peter McLeod, Tim Shallice e David C. Plaut (2000) https://doi.org/10.1016/S0010-0277(99)00067-0