Predictive coding

The general framework Seth imports into the interoceptive domain. Originating with von Helmholtz’s “unconscious inference” and reaching modern prominence via Friston’s free-energy principle and the Bayesian-brain hypothesis (Clark, Hohwy, Friston). See anil-seth.

The primary source, read at last (Friston 2005)

For its first fourteen ingests this page cited “Friston’s free-energy principle” without reading Friston. “A theory of cortical responses” (2005) is the statement of it, and pins down which of this page’s habitual claims are theorems and which are downstream interoceptive extensions. Three things it establishes at the source:

  1. Inference and learning are one optimization. Minimizing prediction error is minimizing free energy, a bound on surprise; perception (the E-step) tunes representations over milliseconds, learning (the M-step) tunes synapses over seconds — the same objective at two timescales. This is the derivation under every “the brain minimizes prediction error” sentence in the wiki.
  2. The architecture is derived, not assumed. Hierarchy, reciprocal forward/backward connections, the forward-drives / backward-modulates asymmetry, and Hebbian plasticity fall out of what empirical-Bayesian inference requires, and match what visual-cortex anatomy already showed.
  3. Evoked responses are prediction error being explained away — a single account of repetition suppression, the MMN and P300, priming and global precedence. See karl-friston.

The paper is exteroceptive throughout (vision, audition); its interoceptive life is entirely Seth’s and Barrett’s later work. That separation is exactly what makes it useful — it shows the machinery undisplaced. The bet it rests on (distinct error and representation unit populations) is the same one flagged as still-open below.

Core mechanics (as summarized in Seth 2013)

  • Perception is not bottom-up feature accumulation; content is specified by top-down predictions from hierarchical generative models of the causes of sensory signals.
  • The brain continuously minimizes prediction error (discrepancy between inputs and model-based predictions) by either (a) updating the model or (b) acting on the world (active-inference).
  • Prediction errors carry precisions (inverse variances). Precision-weighting — plausibly via post-synaptic gain modulation, and interpreted as attention — sets the relative influence of prediction errors vs prior expectations. Low precision on error signals = high confidence in priors.
  • Hierarchical (“empirical”) Bayes: posteriors at one level become priors for the level below, so priors are induced from the data stream itself.

Why it matters here

Predictive coding had been developed almost entirely for exteroception (vision, audition) and action. Seth’s contribution is to note that “one of the most relevant features of the world for an organism is the organism itself” and extend the machinery inward, yielding interoceptive-inference and a predictive account of embodied-selfhood. Precision-weighting becomes the lever explaining how the same visceral arousal can be attended-to (revising feelings) or attenuated (enslaving autonomic reflexes).

Barrett’s version: predictions are concepts

Barrett (2017) imports the same machinery for a different purpose and adds three things the wiki’s Seth-derived account does not have.

1. An identification. “Predictions are concepts. Completed predictions are categorizations.” The competing simulations the brain assembles to answer “what is this new sensory input most similar to?” are the conceptual system — there is no separate store of concepts consulted by perception. See situated-conceptualization.

2. Prediction error processing is concept learning. The upstream sweep is not just error correction; it is abstraction. Error flows from the upper layers of granular sensory cortex (many small pyramidal cells, few connections) toward less granular heteromodal and limbic cortex (fewer, larger cells, many connections), compressing and losing dimensionality as it goes (Finlay & Uchiyama, 2015). “All new learning is concept learning, because the brain is condensing redundant firing patterns into more efficient multimodal summaries.” Dimension reduction is what a concept is made of — and because “conceptually similar representations reuse neural populations,” different predictions end up separable but not spatially separate, laid out in a continuous neural territory organized by similarity. That is a direct prediction about why category-specific voxel patterns overlap. See degeneracy.

3. A why. Seth’s account explains what predictive coding does; Barrett’s explains why a brain would run it at all. “A brain implements an internal model of the world with concepts because it is metabolically efficient to do so.” The internal model costs 20% of the body’s energy (Raichle, 2010) and long-range connections are its priciest component, so compression is not an elegant design choice — it is the budget. See allostasis.

The ordering claim, and what it costs appraisal theory

The consequence Barrett presses hardest, and the wiki’s other predictive sources do not state it this baldly:

In predictive coding, as we will see, sensory predictions arise from motor predictions… For a given event, perception follows (and is dependent on) action, not the other way around.

If that is right, the stimulus → response architecture is not merely incomplete but has no slot for an unevaluated stimulus to occupy. She uses this to reject causal appraisal theories outright — “meaning does not trigger action, but results from it.” See are-appraisals-causes-or-descriptions.

Worth noting a historical hedge she includes and the wiki should keep: the idea that the mind drives perception is old — Ibn al-Haytham in the 11th century, Kant in the 18th, Helmholtz in the 19th, Craik’s internal models (1943), Tolman’s cognitive maps (1948), Neisser, Gregory. What is new in recent formulations is not prediction but four specific commitments: that predictions are embodied simulations, that they are ultimately in service of allostasis (hence interoception at their core), the breadth of anatomical evidence, and the specific computational implementation. “Feedback” itself, she notes, is a term inherited “from a time when the brain was thought to be largely stimulus driven.”

The social destination: prediction error as a motivational cost (Theriault et al. 2021)

Barrett’s version above supplies the why — a brain runs predictive coding because it is metabolically efficient. Theriault, Young & Barrett (2021) make that efficiency claim do motivational work. Their premise: because predictive processing minimizes cost by transmitting only unpredicted signals, the thing a predictive brain pays for is prediction error — so an unpredictable environment is metabolically expensive, and a predictive organism is motivated to make its environment predictable. See metabolic-cost-of-prediction-error.

The payoff is social. Other people are the largest unpredictable part of a human’s environment, so the cost of prediction error becomes a reason to conform to them (the sense-of-should). This is a use of predictive coding the wiki’s other sources do not make: Seth explains perception with it, Barrett explains emotion and concepts with it, Khalsa explains psychopathology with it — Theriault et al. explain motivation with the cost of the error term itself. The paper also engages the standard “dark room” objection (a brain that only minimized prediction error would seek a dark quiet room; Friston et al. 2012) by insisting the goal is never to minimize error but to trade constructing (paying to learn) against coasting (exploiting the model) — the same explore/exploit tension, framed in energetic terms.

A demotion worth recording: one algorithm among several (Petzschner et al. 2021)

Every source above treats predictive coding as the framework. Petzschner et al. (2021) treat it as a hypothesis at a particular explanatory level, and the distinction changes how confidently several claims on this page should be read.

Their separation is Marr’s:

  • Computational levelwhat is computed. The claim that the brain combines a noisy afferent likelihood with a prior, weighted by the precision of each, to estimate a hidden bodily state. This is interoceptive Bayesian inference, and it functions as an ideal-observer benchmark: a standard against which real behaviour is compared, with substantial evidence that humans behave close to it in multisensory integration, in the integration of past experience, and even under abstract beliefs (placebo).
  • Algorithmic levelhow. Bayesian inference “does not come with a prescription for how the computations are implemented,” and “suggestions abound for how the CNS could approximate Bayesian inference with neuronal algorithms.” Predictive coding is “one of the most prominent” — not the only one. The alternatives they cite include probabilistic population codes (Ma et al.), sparse coding (Olshausen & Field), and one flat dissent (Brette, “Is coding a relevant metaphor for the brain?”).

The status report on the algorithmic claim is the sentence to keep: “To date, the full interoceptive brain network underlying this implementation has not been identified.”

So the wiki has been running two claims together. The Bayesian one is well supported and largely orthogonal to which neurons do what; the prediction-error-unit one is an anatomical bet that is elegant, motivated by the agranular visceromotor cytoarchitecture, and still unconfirmed. feedforward-vs-predictive-interoception already records the dispute; what is new here is that the dispute has more than two positions, and that a defender of the Bayesian claim need not defend the coding scheme at all.

Two more consequences of the same framing:

  • Priors replace set-points. In the reflex arc there is a set-point and a comparator; in the inferential account, the states an organism expects to occupy constitute the prior, and they coincide with survivable states because evolution and experience put them there. Whether these are innate or learned is open (“an intriguing open question that might be resolved by developmental studies”) — the wiki’s version of that question is social-vs-biological-origins-of-interoception.
  • The internal/external distinction moves. Interoception and exteroception are distinguished “based on the type of state that is being inferred — internal or external — rather than the specific sensory channels that contributed to the inference.” Vision informing an internal state is interoceptive input. See sensory-control-loop.

The clinical destination: computational psychiatry

The Khalsa et al. (2018) roadmap takes the same machinery in a different direction from Seth’s and Barrett’s — toward psychiatry. If interoception is hierarchical Bayesian inference and control, then interoceptive disorders are specific failures of it: mis-set priors, aberrant prediction errors, or maladaptive precision. The roadmap’s mechanistically loaded borrowing is that belief precision sets the force of regulatory action, which turns pathologically precise body-state priors into over-vigorous regulation — the seed of a “computational psychosomatics.” This is the framework’s clinical payoff and, per the authors, its least-tested part (“indirect so far”). See computational-psychiatry and interoceptive-psychopathology.