Interoceptive sensitivity
The individual-differences construct linking interoception to emotion and selfhood. Operationalized by the heartbeat-detection-task (Schandry 1981). In Seth (2013) it is the trait through which interoceptive precision is probed.
Findings assembled in Seth (2013)
- Right AIC activation and morphometry predict interoceptive sensitivity (Critchley et al. 2004), which in turn predicts emotional symptoms and susceptibility to anxiety.
- Lower interoceptive sensitivity → greater susceptibility to the rubber-hand-illusion and to malleable self–other boundaries (Tsakiris et al. 2011) — interpreted as lower precision-weighting of interoceptive prediction errors, so exteroceptive cues dominate self-model updating. Caveat, from Iodice et al. (2019): this moderation does not generalize to a self-report MAIA measure — MAIA Noticing/Body Listening did not predict the strength of their illusion of effort — which is one more instance of the accuracy(Tsakiris’s objective task)/sensibility(MAIA) non-correspondence catalogued throughout this page. Objective accuracy against that illusion is untested.
- Cardiac timing modulates memory for words as a function of interoceptive sensitivity and metacognition (Garfinkel et al. 2013).
Interpretive note
Within interoceptive-inference, interoceptive sensitivity is not merely “how good your visceral sense organ is” but reflects the precision the system assigns to interoceptive prediction errors — a Bayesian-weight reading rather than a pure signal-fidelity reading. See the terminology caution in the frontmatter: modern usage (Garfinkel’s 2x2) distinguishes accuracy vs sensibility vs awareness, distinctions this 2013-vintage usage predates.
A further wrinkle: sensitivity/attention ≠ accuracy
Farb et al. (2015) highlight a finding that complicates any simple reading of “interoceptive sensitivity”: experienced meditators, despite strong interoceptive attention tendencies, do not show superior heartbeat-detection accuracy (Khalsa et al. 2008; Parkin et al. 2013). This is a central motivation for their expanded interoceptive-taxonomy, and is tracked as an open empirical question at does-mindfulness-enhance-interoceptive-accuracy.
Two more complications from the developmental literature
Oldroyd et al. (2019) add pressure on the construct from a different direction:
More sensitivity is not better. Their attachment-anxious profile scores higher on noticing and emotional awareness while scoring far lower on not-worrying — heightened sensitivity in the service of threat-monitoring. Whatever “interoceptive sensitivity” measures, it is not a quantity of which more is reliably good; the anxious and avoidant profiles are qualitatively different rather than points on one axis. This is the same lesson as the meditator finding above, arriving from developmental psychology.
Part of the trait may be cardiodynamic, not perceptual. Stroke volume predicts heartbeat-detection performance (Schandry et al. 1993), and hpa-axis activation raises stroke volume — so measured “sensitivity” partly reflects how loud the signal is, not how well it is heard. See heartbeat-detection-task. This is orthogonal to the precision reading above but pulls in the same direction: the trait is less about the visceral sense organ than the label suggests.
A third developmental axis: the construct declines with age — but mind the sub-construct
MacCormack et al. (2021) add the downward end of the lifespan (the Oldroyd section above is about its origins), and doing so surfaces the taxonomy problem this page keeps hitting. The claim MacCormack et al. build on — and cite for — is that objective interoceptive ability falls with age (Khalsa, Rudrauf & Tranel 2009; Murphy et al. 2018): older adults do worse on heartbeat tasks. That claim is now read first-hand, and it holds: in 59 adults aged 22–63, age accounted for 30% of the variance in heartbeat-discrimination accuracy, the sole significant demographic (BMI and sex null), reliable across two sessions. So age is a genuine moderator of the accuracy construct this page is named for — the wiki’s first first-hand evidence that its central measure has a large developmental gradient.
But MacCormack et al.’s own studies do not measure accuracy. They measure interoceptive knowledge — how strongly a person’s emotion concepts are associated with bodily sensations (Study 1) — and self-reported sensation intensity (Study 2). Both decline with age (specifically for high-arousal sensations), but neither is the heartbeat-detection accuracy the page’s definition invokes. So the aging finding is a warning as much as a data point: “interoception declines with age” is now confirmed first-hand across three different constructs (knowledge, self-reported intensity via MacCormack; objective accuracy via Khalsa) that are not the same thing and need not have declined together — they did, which is the convergence maturational-dualism wants, but the interoceptive-taxonomy’s recurring complaint still bites: only the accuracy leg is the construct this page is named for, and even that leg is confoundable with cardiodynamics (an older heart is a quieter signal; Khalsa did not measure stroke volume), so its age decline need not be a decline in perceiving. See age-related-interoceptive-decline, maturational-dualism.
A forensic correlate, and what it does and does not add
Nentjes et al. (2013) supply the construct’s first forensic correlate: in 75 offenders, heartbeat-discrimination d’ was inversely related to the antisocial dimension of psychopathy (Factor 2, r=−.29; antisocial Facet 4, r=−.24) and unrelated to its affective/interpersonal dimension (r=−.01). Two things for this page:
- It is a rare direct accuracy↔outcome zero-order correlation — the kind of main effect the Dunn gain-term reading says accuracy does not have. Soft tension only: the “outcome” is a personality trait rather than felt arousal or decision quality, and the effect is modest. It is now the third such accuracy↔outcome main effect the wiki holds first-hand — Pollatos’s felt arousal (r=.34), Nentjes’s antisocial psychopathy (r=−.29), and Werner et al.’s (2009) IGT choice (r=±.30) — against Dunn’s single null (r=.08). The pattern the pure gain-term view does not predict is really a pattern about design: all three main-effect findings come from extreme-groups or clinical samples, and Dunn’s null is the only unselected continuous one, so what accumulates may be selection inflation / mixture purification rather than a main effect the gain-term view has to explain away. See is-more-interoceptive-awareness-better.
- The affective-vs-antisocial split is informative about what the score tracks. If “interoceptive sensitivity” were a straightforward index of emotional depth, it should have loaded on psychopathy’s affective facet; it loaded on the behavioural one instead. That is another way the single label misleads — the construct here indexes something closer to behavioural regulation than to felt emotion. But the reading is fragile: the sample discriminated at chance on average (mean d’=0.00), so this is variation around zero, and an unexplained negative IQ→d’ effect in the same study warns that the residual variance may be partly non-perceptual. See nentjes-2013-psychopathy-interoception, is-more-interoceptive-awareness-better.
The strongest complication: the trait may not be a trait of anything
Dunn et al. (2010) measured this construct with the heartbeat-detection-task in 58 and then 92 participants, and it predicted nothing:
| interoceptive accuracy ↔ | result |
|---|---|
| felt arousal (Study 1) | ns |
| felt valence (Study 1) | ns |
| cardiac response to emotional images (Study 1) | ns |
| intuitive decision quality (Study 2) | r = .08, p = .46 |
| how well one’s own body differentiated good from bad options (Study 2) | r = .07, p = .56 |
Every effect in the paper is an interaction. Interoceptive accuracy determined how tightly bodily responses coupled to felt arousal (ΔR² = .12) and to decision quality (ΔR² = .08) — helping when the body was right, hurting when it was wrong.
The reading that follows is worth stating plainly, because it recasts everything above. Interoceptive sensitivity is not a quantity of a good thing; it is a gain term. It sets how loudly the body’s signal arrives and has no bearing on whether the signal is worth hearing — it does not even correlate with the signal’s usefulness (r = .07). This is the same lesson as the meditator finding and the attachment-anxious profile, arriving for a third time from a third direction, but with a sharper form: those two say more is not reliably better, while this one says the construct has no valence at all and the question “is more better?” is asking a multiplier to behave like an addend. See is-more-interoceptive-awareness-better.
It also gives the precision reading above an empirical shape. If sensitivity is the precision the system assigns interoceptive prediction errors (interoceptive-inference) rather than signal fidelity, then a gain term is exactly what it should look like from the outside — precision-weighting is definitionally a multiplier on prediction error, not a source of accuracy. Dunn et al. do not make this argument (their framing is Jamesian, not Bayesian), and the wiki should not claim they do. But the two vocabularies are describing the same functional role, and that is worth recording.
Two brakes. The result is two unreplicated interactions from one 2010 paper, and no simple-slope tests were run, so whether poor perceivers are uncoupled from their bodies or inversely coupled is undetermined. And the cardiodynamic confound below is not merely present but load-bearing here: if stroke volume raises both the score and the size of the cardiac signal available to couple with, the interaction is what a pure signal-strength account predicts with no perception in it.
The complication to the complication: an earlier study found the main effect
The section above is the strongest claim on this page, and it needs a brake that the Dunn ingest could not supply, because the relevant paper had not been read yet.
Pollatos, Kirsch & Schandry (2005) ran nearly Dunn’s Study 1 five years earlier — affective pictures, SAM arousal and valence ratings, Schandry counting, 44 healthy adults — and interoceptive accuracy did predict felt arousal, as a main effect: F(1, 39) = 5.90, P < .05 (good perceivers 4.83 vs poor 4.19), r = 0.34. Emotion-specifically, too: higher arousal ratings for pleasant and unpleasant pictures, nothing for neutral ones (F(1, 39) = 0.09).
So “it predicted nothing” is a claim about one study, not about the construct. Pollatos et al.’s own introduction records that the field’s consensus ran the other way — Schandry (1981), Wiens et al. (2000), Critchley et al. (2004), Ferguson & Katkin (1996), Montoya et al. (1993) all positive, Blascovich et al. (1992) alone against. The wiki met that consensus only through Dunn’s rebuttal of it, and mistook a rebuttal for a verdict.
And one member of that consensus is now read first-hand, with a better design than Pollatos’s. Wiens, Mezzacappa & Katkin (2000) found the same main effect five years earlier — good heartbeat detectors felt emotion more intensely across amusement, anger and fear (F(1, 50) = 7.61, p < .01), holding continuously (ΔR² = .17, p < .02) — and did the one thing this page keeps faulting the consensus for skipping: it measured bodily arousal and controlled for it. Good and poor detectors did not differ in skin conductance or heart rate, and the intensity effect survived electrodermal activity as a covariate. So the “louder heart → more arousal → more intense feeling” shortcut is closed here (though the stroke-volume signal-amplitude confound is not — SCR is sympathetic outflow, not beat loudness). Wiens also used the discrimination task rather than Schandry counting, so the main effect is not counting-specific. On the arousal main effect the tally is now 2–1 (Wiens + Pollatos find it, Dunn does not) — but Wiens, like Pollatos, never measured a cardiac response to the stimuli, so it cannot test Dunn’s gain-term model, only its margin. The verdict “it predicts nothing” was always one study’s.
Two reasons the table above still stands, and one reason it does not stand as strongly.
It stands because Pollatos et al.’s sample was constructed: 22 good perceivers picked from a screen of ~140, plus 22 matched comparisons. Extreme-groups designs detect group differences more easily than unselected samples and inflate correlations computed across them — so r = 0.34 and r = .08 are not rival estimates of one number. It also stands because Pollatos et al. never measured a bodily response to the pictures (ECG was recorded and used only to score the counting task), so they cannot test the gain-term model at all; they test a model Dunn argues is the wrong one.
It stands less strongly because a main effect is arguably what the gain-term model predicts on the margin. If feeling is b1*(body) + b3*(body x accuracy) with no accuracy main effect, then averaging over bodily responses that are not centred on zero leaves an accuracy slope of b3 x (mean body). Pollatos et al. measured only that margin. This is the wiki’s arithmetic, not either paper’s, and it is incomplete: Dunn’s own marginal was null and the account owes an explanation of why. See pollatos-2005-interoceptive-awareness-erp.
Where the two agree is the durable part. Neither finds any relationship between interoceptive accuracy and felt valence — Pollatos F(1, 39) = 0.14 and no interaction; Dunn delta-R-squared = .00, p = .98. Two labs, two designs, two samples, five years apart. Whatever this trait is, it is not a trait about pleasantness. See core-affect.
And note the confound that survives both: a louder heart is easier to count and a bigger signal to couple to and plausibly comes with higher tonic arousal. That single non-perceptual account predicts Pollatos’s main effect and Dunn’s interaction alike — and predicts Pollatos’s P300 difference being emotion-general (present even for household objects), which no perception-of-emotion account does. The cardiodynamic worry is not one complication among several on this page; it is the one that keeps regenerating.
The trait may not be a trait, and most people may not have it at all
The page’s title, definition and frontmatter all call this a characterological trait. Van der Does et al. (2000) pooled 709 participants across seven studies and put pressure on both words.
Most people do not have the ability. Under a <10%-error criterion, 7.9% of normal controls and 17.1% of panic patients qualify as accurate perceivers. More than 95% of the sample produced a count; roughly 80% were wrong by ~30%. Whatever this page’s construct is, in a typical healthy sample it is a property of fewer than one participant in ten — and the other nine still generate a score.
The ability is not stable, and the inability is. Of 17 panic patients accurate before treatment, 8 were still accurate after; 4 became inaccurate. Prior work had concluded from % error scores that good heartbeat perception is a stable individual characteristic (Ehlers & Breuer 1996; Antony et al. 1994); the categorical re-scoring reverses it. By contrast only 5 of 51 baseline-inaccurate perceivers ever became accurate. The trait-like half of this construct is the absence of the ability — which is what you would expect if most people are not perceiving their hearts at all, and equally what you would expect if the trait is cardiodynamic and treatment does not change stroke volume.
And it can be manufactured. Exercise transiently produces accurate perceivers above ~100 bpm; by ~95 bpm they are gone. One of 60 participants improved durably. This is the cardiodynamic complication two sections above, no longer a correlation but a manipulation — and it is hard to call something a characterological trait when a treadmill moves it and rest moves it back.
What it does to the section above. The “gain term” reading says accuracy is a multiplier on the bodily signal. The mixture reading says something more awkward: for ~80–90% of participants there may be no coefficient to estimate, because the measured variable is not tracking the construct. A correlation computed across a mixture of one valid subpopulation and one measuring-something-else subpopulation is attenuated toward zero — which is the shape of r = .08 and r = .07. So Dunn’s nulls have a candidate explanation that is neither “the construct has no valence” nor “the effect is real but small,” but “the moderator is mostly noise in an unselected sample.” That does not rescue the main effect; it means the wiki’s three explanations for the Dunn/Pollatos disagreement (sampling inflation, missing bodily measurement, mixture attenuation) are not mutually exclusive and the data cannot separate them. See is-the-heartbeat-counting-task-valid.
Part of what it does measure is a belief. Accurate perceivers scored ~half an SD higher on anxiety sensitivity (26.9 vs 21.0) and differed on nothing else — not trait anxiety, not state anxiety, not depressive symptoms, not somatosensory amplification, not age, BMI, or actual heartbeat count. In interoceptive-taxonomy terms, the wiki’s canonical accuracy measure is partly indexing a sensibility construct, which is the taxonomy’s own worst-case scenario.
Two brakes on all of this. The categorical scoring the argument depends on rests on a bimodality claim never formally tested (three histograms, no dip test, no mixture model — see van-der-does-2000-heartbeat-perception-reanalysis). And it is a clinical-anxiety sample, cardiac only; whether the minority-validity thesis generalizes to respiratory or gastric interoception is untouched.
The complication that subsumes several of the others: it is not a trait of the person at all, but of an organ
Every section above argues about what this construct is a trait of — a perceptual ability, a gain term, a cardiodynamic fact, a belief. Banellis et al. (2026) make a prior point: whatever it is, it is not a property the person carries from one organ to another.
In 241 participants measured on cardiac (HRDT) and respiratory (RRST) axes with a single hierarchical Bayesian pipeline, cardiac and respiratory sensitivity correlated r = −0.019 (BF01 = 6.39). Same for precision and for metacognitive efficiency. The design had 80% power at r ≥ 0.179.
This page’s title, and its definition — “a characterological trait reflecting individual sensitivity to interoceptive signals” — assert exactly what the study fails to find. Two of the three words are now in trouble: Van der Does took “trait,” and this takes “interoceptive.” What survives is cardiac sensitivity, measured in a channel, generalizing to nothing tested.
What it does to the sections above. Mostly it narrows their scope rather than refuting them. The gain-term reading, the cardiodynamic confound, the anxiety-sensitivity contamination and the aging gradient are all claims about the cardiac measure, and all stand as such — indeed the cardiodynamic confound is strengthened, since a channel-specific signal-amplitude property is precisely what would produce a channel-specific “trait.” What no longer follows is the inference each of those sections quietly makes at its end: from a fact about heartbeat scores to a claim about interoception.
One thing did generalize, and it is the wrong one. Mean confidence correlated across cardiac, respiratory and auditory tasks (r = 0.51–0.64). So the part of this construct that behaves like a person-level trait is the subjective part — which is the sensibility side the whole accuracy/sensibility distinction was built to quarantine. See metacognitive-efficiency, is-interoception-domain-general.
And the word “trait” fails on a second axis: time
Banellis et al. above take “interoceptive” out of the definition by showing the construct does not travel between organs. Wallman-Jones et al. (2023) press on “trait” from the direction Van der Does opened, and with a design built for it: ten prompts a day for seven days in 70 people yields an ICC of 0.51 for self-reported interoception, and the state moves with ordinary behaviour — up with movement, down with sitting and screens.
This is sensibility, not the accuracy this page is named for, and that limit matters. But three of the sections above are already about instability in the accuracy measure — treatment abolishes it, exercise manufactures it, state anxiety moves counting error within a person — and the wiki has read all of them as evidence that the task is confounded. A cleaner reading is now available: the confound story and the state story make the same prediction, and no data on this page separates them. What follows either way is arithmetic, and it is the section this page most needs: a measure with an ICC near 0.5 estimated from one session attenuates every correlation it enters. The Dunn dispute is over r values of .08 to .34, all from single sessions, and nobody in it has corrected for that. See state-vs-trait-interoception.
One result cutting the other way, recorded because it is inconvenient. In the same study, baseline heartbeat-counting accuracy positively predicted week-averaged self-reported interoception (B = 4.50, β = 0.21, p = .007) — the accuracy↔sensibility association this page and maia both say should not be there. Held loosely (N = 70, CI 1.31–7.69, a state instrument rather than a trait questionnaire), but not filed away.
And the construct on the other side of the dissociation is also two things
Nearly every section above is about what happens when this page’s construct — measured accuracy — fails to behave. The comparison class throughout is sensibility, treated as a single alternative: the questionnaire measure that keeps failing to correlate with accuracy, and that Harrison et al. found carrying more of the affective variance.
Ventura-Bort et al. (2021) show that comparison class is itself two orthogonal factors — believed accuracy and deployed attention — with opposite relations to well-being and to emotional granularity. See interoceptive-taxonomy for the detail.
What that does to this page is small but worth stating. The accuracy↔sensibility dissociation recorded in several sections above is a dissociation against a mixture, and nobody has estimated it separately against the two halves. If measured accuracy relates to one factor and not the other — the Wallman-Jones positive association above (β = 0.21) used a state attention-flavoured instrument, so the question is live — then the pattern of nulls this page treats as one finding may be two findings averaged. Untested: Ventura-Bort et al. administered no accuracy task at all.
And when someone did run the accuracy task, the nulls turned out to be about attention
The paragraph above names the missing study. Murphy et al. (2020) had already run it — the IAS Ventura-Bort et al. borrowed was built for exactly this test, and it was published alongside a heartbeat task.
The result is the one this page has to absorb. In two samples, heartbeat-counting accuracy correlated with a self-report of believed accuracy (r = .271, n = 52; r = .336, n = 33) and with nothing else administered — not the BPQ, not the ICQ, not the TAS-20, not time-estimation performance. Post-task confidence, the state form of believed accuracy, behaved identically.
So the dissociation this page is largely organized around has a boundary, and it runs along the what axis. Go back through the sections above and check what sat in the self-report slot each time: maia Noticing (Iodice), MAIA Noticing again (Oldroyd), MAIA and BPQ (Harrison), a state mindfulness scale (Wallman-Jones). Attention instruments throughout. Match the construct being asked about and the correlation appears; mismatch it and it does not, in the same people, in the same session.
Two consequences for how this page should be read.
The Wallman-Jones positive result stops being an anomaly. It is filed above as “one result cutting the other way, recorded because it is inconvenient.” It now has company and a mechanism: accuracy relates to the believed-accuracy half of self-report. Two positive findings on different instruments is a pattern, not a pair of flukes — though Wallman-Jones’s instrument is attention-flavoured, which the account above predicts should have been null, so the convergence is thematic rather than exact.
It does not rescue the construct, only the vocabulary. Every objection on this page survives untouched: the correlation is with a counting score in an unselected sample, so cardiodynamics, the mixture problem and the state instability all apply to the criterion as much as ever. And r ≈ .3 at n = 33–52, p ≈ .047, is not a large or a secure number. What changes is narrower and worth stating exactly: the repeated failure of “sensibility” to track accuracy is a failure of attention measures to track accuracy, which is what the taxonomies predicted all along and which nobody had tested with the what axis held fixed. See murphy-2020-interoceptive-accuracy-scale.
A third complication: the trait may be pharmacologically movable
Lyons et al. (2021) relay a review finding the wiki should hold onto: interoceptive accuracy in depression is not simply low but inconsistent, pointing toward low IA in moderately depressed and better IA in severely depressed individuals (Eggart et al. 2019). Comorbid anxiety is one proposed confound — anxious patients monitor the body closely (Paulus & Stein 2010) — and studies controlling for it do find reduced IA in depression (Furman et al. 2013).
The part with teeth: Eggart et al. additionally flag antidepressant medication as a possible confound on IA. If a drug taken by most patients in most depression samples moves the measure, then a large fraction of the clinical interoceptive literature is confounded by treatment status, and “interoceptive accuracy in depression” is partly a fact about pharmacology. Lyons et al. supply a suggestive parallel on a different instrument: medicated patients’ bodily-sensation-maps were dominated by deactivation in a way unmedicated patients’ were not. See antidepressant-emotional-blunting.
Caveat: none of this is measured here. Lyons et al. ran no heartbeat task, so the medication/IA link is inherited by citation and the medication/BSM link is correlational and exploratory. Recorded as a confound to check, not an established one — and as one more reason the single “sensitivity” label is doing too much work (see interoceptive-taxonomy, is-more-interoceptive-awareness-better).