A theory of cortical responses
The primary source for machinery this wiki has been running secondhand since its first ingest. Every predictive-coding page here — predictive-coding, active-inference, perceptual-inference, interoceptive-inference — traces back to “Friston’s free-energy principle” without ever reading Friston. This is the paper that states it: the 2005 Phil. Trans. R. Soc. B article where free energy, empirical Bayes, hierarchical generative models, and prediction-error/representation unit populations are assembled into one account of what a cortical response is. See karl-friston.
It is not an interoception paper. Its worked examples are vision (V1/V2 receptive fields, extra-classical surround effects, repetition suppression of face responses) and audition (the MMN). Its value to this collection is exactly that it is the undisplaced statement of the framework — so the wiki can see which of its habitual claims are load-bearing theorems and which are Seth’s and Barrett’s later interoceptive extensions.
The one identity the whole framework rests on
Both inference and learning rest on minimizing the brain’s free energy, as defined in statistical physics.
Friston’s move is to notice that two things neuroscience treats separately — perceiving (inferring the causes of sensation) and learning (tuning the model that relates cause to sensation) — are the same optimization at two speeds. Formally they are the E-step and M-step of expectation maximization on one objective function, the free energy F, which is a lower bound on the surprise (negative log evidence) of the sensory data. The E-step (inference) adjusts the conditional expectation of causes, φ, over milliseconds; the M-step (learning) adjusts the parameters θ (synaptic strengths) over seconds and longer. “Maximizing the objective function (i.e. minimizing the free energy) is simply minimizing our surprise about the data.”
This is the sentence under every downstream use. When the wiki says the brain “minimizes prediction error,” this is the derivation being invoked.
Why hierarchy and backward connections are forced, not assumed
The paper’s rhetorical strategy is to derive cortical anatomy from the statistics of inference, then check it against what neuroanatomy already knows. Because sensory causes mix nonlinearly (occlusion, context), the generative model that produces sensation cannot be simply inverted — recognition is ill-posed and needs priors. Empirical Bayes supplies them by treating a hierarchical model’s higher levels as priors on its lower levels, so the priors are learned from the data stream itself rather than built in. The implementation then requires:
- Hierarchical organization — levels whose causes are elicited by the level above.
- Reciprocal connections — forward error-carrying and backward prediction-carrying projections between each pair of levels.
- A functional asymmetry — forward connections drive (they always elicit a response), backward connections both drive and modulate (they can be nonlinear, are more divergent, and act through slow NMDA channels in supragranular layers). Friston is careful here: he calls forward connections the feedback connections, because in a generative model the world’s causal structure lives in the backward projections and the forward stream is just the error correcting them.
- Associative plasticity — the M-step is Hebbian: connection strengths change in proportion to pre- and post-synaptic activity.
All four were already documented in visual cortex. The claim is that they are what inference requires, not incidental wiring.
Evoked responses are prediction error being explained away
The step that matters most for the rest of neuroscience: a stimulus-evoked response can be decomposed into two subpopulations — representational units encoding the conditional expectation of causes, and error units encoding prediction error. Because higher levels take time to settle on the cause, the error signal at each level waxes and then is suppressed as the explanation arrives. This makes an evoked transient look like a damped oscillation, and predicts that late components of ERPs reflect inference about higher-order, more global causes — which is what the empirical literature shows (selectivity for stimulus identity and expression emerging after the initial visual response; the N170; the posterior N2 enlarged by global/local incongruity).
Two heuristics from the vision literature — high-level areas “explain away” prediction error, or high-level areas tell lower areas to “shut up / stop gossiping” — are shown to be the same thing under empirical Bayes: predictions explain away error while lateral interactions among error units select the causes. The apparent tension dissolves into the two-subpopulation architecture.
The MMN, and the link to ERPs
The paper’s cleanest empirical anchor, and the reason it is filed against the wiki’s ERP method page. The mismatch negativity — a negative ERP component to a deviant in a train of standards, seen without attention — is recast as the failure to suppress prediction error when a stimulus violates the statistical regularity the system has learned. On this reading the MMN is not generated by dedicated “change-detection neurons”; it is what prediction error looks like before learning-related plasticity (the M-step, in backward and lateral connections) attenuates it over repeated exposure. Two consequences the paper presses:
- Repetition suppression (the MMN shrinking as standards accumulate, “roving” paradigm) is the M-step happening in real experimental time.
- Pharmacology fits. Ketamine, an NMDA antagonist, cuts the MMN ~20% (Umbricht et al. 2000). NMDA receptors carry the slow plasticity the M-step needs, so blocking them should compromise exactly this suppression. Friston ties this to his disconnection hypothesis of schizophrenia — aberrant experience-dependent plasticity — making the MMN a candidate quantitative index of perceptual learning in psychiatric research.
This is a stronger, mechanistic claim than the wiki’s existing ERP content, which treats components as non-specific indices (“the P300 indexes attention, capacity, relevance and difficulty at once”). Friston’s account says what the component is — suppressed prediction error — at the cost of committing to the predictive-coding architecture to say it. The event-related-potentials page’s standing caution (“trust the latency, discount the anatomy”) applies with full force: the MMN-as-prediction-error reading is a theoretical interpretation of scalp timing, not a source-localized measurement.
The empirical epilogue: connectivity changes carry the MMN
To show the framework is not merely a re-description, the paper ends with a dynamic causal model (DCM) of an auditory-oddball ERP dataset. Six sources (bilateral A1, orbitofrontal, superior temporal, posterior cingulate) are modelled with neural-mass dynamics; Bayesian model selection compares four DCMs allowing standard-vs-deviant differences in forward-only, backward-only, forward+backward, and forward+backward+lateral connections. The forward+backward+lateral model wins by a log-evidence margin of 27.9 (“very strong evidence”). The theory predicted that perceptual learning must involve backward and lateral plasticity; the data require exactly those connections to change. The honest limits are stated in the paper: no functional roles are assigned to the three cell populations, and the sign of the connection changes is left uninterpreted.
Where this sits relative to the wiki’s interoceptive uses
- Seth (2013) imports this machinery inward: the same top-down prediction / bottom-up error / precision-weighting scheme, but over interoceptive rather than visual causes, with the AIC as comparator and autonomic outflow as active-inference.
- Seth & Friston (2016) is the direct interoceptive sequel co-authored by Friston himself, adding the cytoarchitectural (agranular) argument and the epistemic/instrumental split of active inference.
- Barrett (2017) uses the same architecture to a different end — predictions as concepts, prediction-error minimization as concept learning, and the whole apparatus justified by the metabolic cost of running it — the “why” this paper does not supply.
- Petzschner et al. (2021) is the necessary counterweight: it treats predictive coding, this paper’s implementation, as one algorithmic hypothesis for a computational-level claim (Bayesian inference) that could be realized other ways. Read Friston 2005 as the strongest statement of the coding scheme, and Petzschner as the reminder that a defender of the Bayesian claim need not defend it.
- Harrison et al. (2021) is the wiki’s first human model-based test of interoceptive predictive coding — and notably it failed to find the anterior/posterior prediction-vs-error dissociation this architecture predicts, a first crack in the empirical picture Friston’s theory paints so cleanly for exteroception.