Is more interoceptive awareness better?

A debate the wiki had been having on at least five pages without a page of its own. It is created here because Lyons et al. (2021) made the collision explicit: two papers using the same instrument, one lineage apart, read reduced felt bodily emotion as benign in one population and pathological in another.

The collision

sourcepopulationfindingverdict
volynets-2020-cultural-universalityageing adults, 18–90felt bodily emotion dampens with agegood — easier regulation, higher life satisfaction
maccormack-2021-aging-emotionsadults 18–75, two studiesemotion–body association weakens with age (high-arousal specific); worse gut decisionsgood for regulation, bad for decisionsmaturational-dualism; same fading signal cuts both ways
lyons-2021-body-maps-depressionmajor depressionfelt bodily emotion is weaker (but see caveat)bad — increase body awareness in treatment
antidepressant-emotional-bluntingmedicated MDDfelt bodily emotion turns to deactivationunwanted — patients complain of numbness
oldroyd-2019-attachment-interoceptionattachment-anxiousheightened noticing and emotional awarenessbad — hypervigilance, low not-worrying
farb-2015-interoception-contemplative-healthanorexia nervosaawareness training raises bodily contactbad initially — increases maladaptive behaviour
farb-2015-interoception-contemplative-healthgeneral/wellbeingcontemplative training raises contactgood — but selectively, never maximally
bechara-2005-somatic-markersVM/amygdala lesion patientssomatic reactivity to loss is absentgood, in one setting — they keep investing where anxious normals quit, and outperform them
dunn-2010-listening-to-your-heart92 healthy adultsaccurate heartbeat perceptionneither — r = .08 with decision quality; helps or hurts only as a function of whether the body was right
werner-2009-cardiac-perception-decision-making50 healthy adults (extreme groups)accurate heartbeat countingmore → better decisions — good perceivers choose more advantageous IGT decks (r=±.30); but net gain ns, decision-quality not wellbeing, and extreme-groups; the main effect Dunn’s row reports as absent
pollatos-2005-interoceptive-awareness-erp44 healthy adultsaccurate heartbeat perceptionmore intense — affective pictures rated more arousing (r = 0.34), bigger P300 and slow wave; no claim about better
wiens-2000-heartbeat-detection-emotion52 healthy adultsaccurate heartbeat discriminationmore intense — films felt more intense across all valences (F = 7.61), arousal-controlled; no valence effect; no claim about better
van-der-does-2000-heartbeat-perception-reanalysis (via Ehlers 1995)panic disorder, 1-year prospectiveaccurate heartbeat perceptionbad — predicts poor treatment outcome and panic recurrence; the page’s only prospective clinical outcome, and it deflates (see below)
nentjes-2013-psychopathy-interoception75 forensic offenders, cross-sectionalaccurate heartbeat discrimination (d’)less is worse — lower accuracy tracks the antisocial factor of psychopathy (r=−.29), not its affective core; a “more is better” datum via the somatic-marker route, but sample mean d’=0.00 (chance) and the affective-deficit hypothesis failed
farb-2010-neural-expression-sadnessMBSR completers vs waitlist, N=36, cross-sectionalinteroceptive (right insula) recruitment during a sad filmmore (interoceptive) is better — trained group keeps insula online where controls deactivate it, and more insula recruitment → lower depression (r=-.465); the page’s one genuinely interoceptive ‘more is better’ point, weakened by a cross-sectional design
banellis-2026-body-wandering536 healthy adults, resting-state fMRIspontaneous body-oriented thought (self-report, not accuracy)worse in the moment, better as a trait — negative affect and high autonomic arousal, yet lower ADHD and depression symptoms; depression instead tracks mental time travel. The page’s only row where the two verdicts diverge within one sample
farb-2011-relapse-prediction / farb-2022-relapse-biomarkersremitted depression, 18-month + 2-year prospective (2022 = RCT, N=85)sensory (visual in 2011; somatosensory/insular in 2022) vs elaborative (prefrontal) reactivity to sad filmswhich mode, not how much — elaborative reactivity predicts relapse, sensory reactivity protects; the sensory pole is largely exteroceptive but moves toward interoception in 2022, and felt intensity predicts nothing

| wallman-jones-2023-everyday-behaviors-interoception | 70 healthy adults, 7-day ambulatory assessment | self-reported interoception after activity vs sitting | more in the moment, less as a habit — momentary activity raises it, habitual activity lowers it (β = −.29), and high-accuracy perceivers decrease where low-accuracy ones increase; no outcome measured, so no verdict on ‘better’ at all |

| ventura-bort-2021-sensibility-conceptualization | 109 healthy adults, 8 self-report instruments | sensibility, factor-analysed into believed accuracy vs deployed attention | both, oppositely — believed accuracy tracks well-being (r = .58) and finer negative granularity; deployed attention tracks worse granularity, lower intensity and r = .12 with well-being. The page’s first row to show the measure summing a good and a bad disposition |

| harrison-2021-breathing-interoception-anxiety | 60 healthy adults, extreme trait-anxiety groups | respiratory detection threshold (FDT) | less is worse — anxious participants detect an added breathing load less well and feel less confident, insight unchanged; opposite in sign to the cardiac panic row |

| eusebio-2022-yoga-interoception-attention | 58 randomized / 34 analysed, sub-clinical symptoms, 10-week RCT | trait sensibility as a moderator of two matched trainings | depends on the training — high sensibility predicts attentional gain under interoceptive instruction and no gain under movement instruction; on the page’s own question (does interoceptive emphasis raise interoception or reduce symptoms more?) the answer is a flat null in both arms |

| rai-2026-dtx-exercise-performance | 70 undergraduates, 3-week activated RCT (digital mindfulness vs cognitive training) | induced change in sensibility vs exercise endurance | more awareness, no payoff — training raised MAIA and slightly raised endurance, but the two were uncorrelated across people (r=−0.02) and awareness mediated nothing; distress tolerance tracked performance instead. ‘Which mode, not how much,’ measured directly |

| murphy-2022-interoceptive-propensity | theoretical commentary | use of interoceptive signals (not perception) | neither — use is a conditional gain term — deploying the body amplifies whatever it signals, adaptive only when the signal is right and insight is good; women’s poorer lab accuracy with better emotion recognition shows external cues sometimes win. No data, no measure, trait/state open |

Read down the “verdict” column and the naive dose-response reading dies. More is sometimes good, sometimes bad; less is sometimes good, sometimes bad. Both extremes have a pathology attached: the anxious hypervigilant end and the numbed/dissociated end.

The row that measured it

The Dunn et al. (2010) row is different in kind from the six above it, and the difference is the reason this page should be reorganized around it rather than merely extended.

Most other rows infer a verdict about interoceptive contact from a population comparison: ageing adults feel less, depressed patients feel less, anxious patients feel more, meditators are trained to feel differently. None of those measured interoceptive ability and an outcome in the same people and asked what the function looks like. Dunn et al. did, in 92 healthy adults, and the function is not a function of interoception at all:

(Correction, with the Pollatos et al. (2005) ingest: this section used to say Dunn et al. were the only source on this page to do that. They are not — Pollatos et al. measured interoceptive accuracy and felt arousal in the same 44 people five years earlier, and reached the opposite conclusion. See “The row that disagrees” below. The claim was written when Dunn was the only such source the wiki had read, and it should not have been future-proofed as strongly as it was.)

  • Interoceptive accuracy ↔ decision quality: r = .08, p = .46.
  • Interoceptive accuracy ↔ how well the person’s own body differentiated good options from bad: r = .07, p = .56.
  • Interoceptive accuracy ↔ felt arousal, felt valence, cardiac response (Study 1, n = 58): all ns.

Then the interaction, which is where everything happens: better perceivers decided better when their anticipatory bodily signals favoured the profitable decks and worse when those signals favoured the unprofitable ones (ΔR² = .08). Same in Study 1 for feeling: better perceivers’ hearts coupled more tightly to their felt arousal (ΔR² = .12) — and not at all to felt valence (ΔR² = .00; the two moderations differ, Z = 2.45, p < .01).

So interoception is a gain term. It sets how loudly the body’s signal arrives, and has no view on whether the signal is true. It does not even correlate with the signal’s truth (r = .07). Amplify a body that is right and you decide better; amplify a body that is wrong and you decide worse; the amplifier is innocent either way.

That reframes both candidate resolutions below rather than choosing between them.

It is bad news for resolution 1 (the inverted U). This page says no source has “measured interoceptive contact and wellbeing across a full range in one sample and looked for curvature.” One now has — for decision quality rather than wellbeing, so the item stands for wellbeing — and found no curve, because it found no line. An inverted U still requires interoception to be on the outcome axis. Here it is on the axis that scales another variable’s effect. You cannot have an optimal dose of a multiplier without specifying what it multiplies.

It is good news for resolution 2, and sharpens it. Resolution 2 says the question is malformed because “interoceptive awareness” is not one thing. Dunn et al. suggest it is malformed for a second, independent reason: even holding the construct fixed at a single well-measured one (accuracy, heartbeat-detection-task), “is more better?” has no answer, because the construct is not the kind of thing that has a valence. The two malformations compound. Sort the rows by construct and ask what each construct multiplies, and most of the collision above is dissolved — Volynets’ quieter body and Farb’s better relationship to the body are not competing doses; they are changes to different terms in a product.

What stops this settling the page. It is one paper, two moderations, n = 58 and n = 92, unreplicated, from the lab that built both the hypothesis and the task, in 2010 — see the limitations on dunn-2010-listening-to-your-heart. It measures decision quality, not wellbeing, and the clinical disagreement at the bottom of this page is about wellbeing. And the moderator may be partly cardiodynamic rather than perceptual (a louder heart is easier to count and a bigger signal to couple to), which would make the interaction a signal-strength effect in perceptual clothing. Recorded as the best-shaped evidence on this page, not as the answer.

The row that disagrees

Pollatos, Kirsch & Schandry (2005) is the second source on this page to measure interoceptive ability and an outcome in the same people — and it found the main effect the section above is built on the absence of.

Nearly the same experiment as Dunn’s Study 1: affective pictures for 6 s, SAM arousal and valence per picture, Schandry counting, healthy adults (44 vs 58). Good heartbeat perceivers rated affective pictures as more arousing — F(1, 39) = 5.90, P < .05; r = 0.34 with the heartbeat score — and the effect was emotion-specific (pleasant and unpleasant yes, neutral no).

Pollatos et al. (2005)Dunn et al. (2010) Study 1
accuracy → felt arousalF(1, 39) = 5.90, P < .05ns
accuracy → felt valencensns
body measured during the task?noyes
sampleextreme groups (22 selected from ~140 + matched)unselected

This does not restore the dose-response reading, and that is the first thing to notice. Pollatos et al.’s verdict column says more intense, not better. They measured how aroused people felt, not whether they were any better off — no wellbeing outcome, no decision quality, no symptom measure. So the row adds a data point about intensity of experience and none about value, which is what this page is actually asking about. On the page’s own question the row is silent, and the sharpest thing it does is complicate the source the page most recently reorganized around.

On the arousal main effect the wiki is now genuinely unsettled. Two considerations favour Pollatos: extreme-groups sampling gives more power to detect a group difference than Dunn’s unselected sample; and a marginal main effect is arithmetically what Dunn’s own interaction predicts once you average over bodily responses that are not centred on zero (this wiki’s reading, and incomplete — Dunn’s own marginal was null). Two favour Dunn: that same selection inflates Pollatos’s r = 0.34 so it is not comparable to Dunn’s r = .08; and Pollatos et al. never measured a bodily response, so they cannot test the gain-term model, only its margin.

What survives intact is the gain-term reading’s shape, not its evidential monopoly. Nothing in Pollatos et al. touches the interaction — they had no bodily-response term to interact with. Their result is compatible with the gain term and does not test it. So the two candidate resolutions below are unaffected: the inverted U still has nothing testing it (Pollatos measured no outcome to be curved against), and the malformed-question resolution still stands.

What does not survive is this page treating Dunn’s null as the last word. Pollatos et al.’s introduction records that the pre-2010 consensus ran the other way — Schandry (1981), Wiens et al. (2000), Critchley et al. (2004), Ferguson & Katkin (1996), Montoya et al. (1993) — with a single dissent (Blascovich et al. 1992). The wiki knew that consensus only through Dunn’s rebuttal of it and adopted the rebuttal wholesale. Two members are now read first-hand — Pollatos and Wiens — and neither goes quietly; Wiens is the one that measured bodily arousal and showed the intensity effect survives it, which is the design the rest of the consensus never ran. On the arousal main effect the balance is now 2-1 against Dunn’s null, and the dissenting design is the least controlled of the three.

The convergence is worth more than the conflict. Both studies find interoceptive accuracy has no relationship to felt valence — Pollatos F(1, 39) = 0.14, Dunn delta-R-squared = .00. Independently, five years apart. See core-affect.

The row that dissolves, and the question it raises about all the others

The Van der Does et al. (2000) row is the only one on this page whose outcome is clinical, prospective, and unambiguous — and it is here mainly as a cautionary tale about the other rows.

The finding. Ehlers (1995): good heartbeat perception predicts poor treatment outcome and recurrence of panic attacks after initial remission, one year out. Not a cross-sectional group difference, not a proxy — an ability measured at baseline predicting a bad clinical outcome later. (Ehlers 1993 now supplies its preliminary version first-hand: 17 remitted patients, the 6 who relapsed having shown better baseline perception — the 1995 paper is the larger replication.) On this page’s actual question (“is more better?”) it is the cleanest “no” available, and the page had been complaining it lacked exactly this: the Pollatos row measures intensity not value, the Dunn row measures card games not wellbeing, the Volynets row is a proposal.

The deflation. Then Van der Does et al. pooled 709 participants and asked what actually distinguishes accurate perceivers. The answer: anxiety-sensitivity, by about half a standard deviation (26.9 vs 21.0), and nothing else. Not trait anxiety. Not state anxiety. Not depressive symptoms. Not somatosensory amplification. Not age, BMI, or actual heart rate during the trials.

So the chain “perceive your heart well → do badly in treatment” plausibly reads “believe bodily sensations are harmful → score higher on a heartbeat task → do badly in treatment,” in which the perception step is a passenger. The authors say so explicitly: higher anxiety sensitivity “may be the reason why good HBP is predictive of worse outcome.”

A within-subject echo of the deflation. Zoellner & Craske (1999) independently replicate the panic-group accuracy advantage (in analogue infrequent panickers) but add no outcome, so they are not a row here. Their contribution to this row is the deflation’s mechanism: within a person, heartbeat-counting error fell as state anxiety rose, which is what the count-inflation artefact predicts and which they instead read as attention. So the same score that predicts worse panic outcome also moves with a person’s momentary anxiety — one more reason to suspect it indexes a relationship to arousal (belief, attention, or inflation) rather than a fixed perceptual ability. See is-the-heartbeat-counting-task-valid.

Why this should worry the whole page. Every row here has the same structure — a measure of interoceptive contact on one side, a verdict on the other. This is the first row where somebody checked what the measure correlates with, and the answer was a belief about the body, not a fact about perceiving it. The Oldroyd row is already half-way to the same place (maia measures sensibility, i.e. beliefs, by construction, and the page says so). The question the Ehlers row leaves behind is how many of the others would survive the same audit — and the honest answer is that nobody has run it for the Volynets, Lyons, or Farb rows.

The measurement worry underneath the Pollatos/Dunn dispute

Van der Does et al. also change the terms of the disagreement two sections above, in a direction this page did not anticipate.

The page’s account of why Pollatos found an arousal main effect and Dunn did not has been sampling: extreme groups inflate correlations and ease detection, so discount Pollatos’s r = 0.34. Van der Does et al. supply a rival reading of the same fact. If the heartbeat task is valid only for the ~8–17% who are genuinely accurate — and the rest produce counts unrelated to their hearts — then Pollatos’s screen of ~140 down to 22 good perceivers (≈16%) is approximately isolating the subpopulation the task works on, and Dunn’s unselected correlation is diluted by ~80% noise, attenuating toward zero.

Selection as purification, not just distortion. Both can be true simultaneously.

What this does to the page. It does not restore the dose-response reading, and it does not touch the gain-term reading’s shape (Van der Does et al. have no bodily-response term either). What it does is add a third live explanation for the wiki’s sharpest empirical disagreement, and make the disagreement harder rather than easier: sampling inflation, missing bodily measurement, and mixture attenuation all predict the observed pattern, and no data on this page separates them. The discriminating test listed at the bottom of this page — measure accuracy, bodily response and felt arousal in one unselected sample — now needs an addition: it must also report the accuracy distribution, because if the mixture reading is right, an unselected sample is the wrong instrument and the test as specified cannot work.

And a caution for both sides: if measured accuracy is state-dependent and manufacturable by raising heart rate (exercise creates accurate perceivers above ~100 bpm and they vanish by ~95), then “good perceiver” is not a stable category to select on or to correlate with. Pollatos selected on one session and analysed as though selecting on a trait. See is-the-heartbeat-counting-task-valid.

The row that comes from inside the house

The Bechara row is new with the Bechara & Damasio (2005) ingest and is worth separating from the others, because of who is reporting it.

Shiv, Loewenstein, Bechara, Damasio & Damasio found that in an investment task where normal individuals drop out because of heightened anxiety after a streak of losses, “the poor somatic reaction of neurological patients to these losses enable them to continue investing, and thus outperform normal individuals.”

Every other row in the table above is a finding about healthy or clinical variation, and most were produced by researchers with a stake in body-awareness being good. This one is produced by the authors of the somatic-marker-hypothesis — the framework built on the claim that these patients’ missing somatic signal is precisely what ruins their decisions — reporting that under a specifiable condition the deficit wins. It is the strongest form the “less is better” position takes anywhere in this wiki: not a benign reduction (Volynets), not a different profile (Oldroyd), but a measurable performance advantage attributable to not feeling the body.

Correction, from the Dunn et al. (2010) ingest: it is not unpublished. This page carried it as “unpublished observations” reported in Bechara & Damasio (2005), and told the wiki to upgrade or drop the claim when Shiv et al. (2005) became available. It had been available all along — Investment behavior and the negative side of emotion, Psychological Science, 16, 435–439 (2005) — cited in full by Dunn et al., who endorse it as agreeing with their own result: absence of emotion after frontal injury “can in some circumstances lead to superior decision making.” The wiki carried a published paper in a major journal as an in-house rumour for three ingests, because the review it first met the finding in described it that way. A lesson worth keeping about inheriting a source’s characterization of its own citations.

Two reasons to still hold it loosely. It remains unread first-hand — not in raw/ — so the n and statistics are unknown here, and this page should not quote effect sizes it has never seen. And the task is an outlier: a rigged investment game in which the optimal strategy is to keep betting through losses. Anxiety-driven exit is a bad strategy there by construction, which is not obviously the world’s usual arrangement.

And Dunn et al. (2010) is the general case of it. Shiv’s patients are a point on Dunn’s surface: bodily signals favouring the wrong action, poorly perceived, hence weakly transmitted, hence better outcomes. Stated that way the “counterexample from inside the house” stops being a paradox and becomes the prediction — which is what a moderation account buys you, and why the row above the Bechara row now does more work than this one.

And one reason it still bites. Bechara & Damasio absorb the result via their integral/unrelated distinction (§2.4): emotion helps when integral to the decision, hurts when unrelated. But the paper offers no criterion for classifying a somatic state as integral or unrelated before seeing whether it helped. Applied here, the distinction is doing no work — it names the outcome rather than predicting it. So the framework’s own escape from its own counterexample is, on this evidence, unavailable. See does-somatic-feedback-guide-decisions.

The row where the verdict splits in two

Banellis et al. (2026) is the newest row and the first to make this page’s question ambiguous within a single sample rather than across them.

Every row above resolves to one verdict because it reports one association. Volynets: less felt body in ageing, and ageing looks fine. Lyons: less felt body in depression, and depression is bad. Dunn: no association at all. Banellis et al. report two, in the same 536 people, pointing opposite ways:

what was correlated with body-wanderingdirectionverdict
momentary affective tone during the scanmore negative, less positive (p_FDR ≤ .001)worse
autonomic arousal (HR up, RMSSD down)higher arousalworse
ADHD symptoms (ASRS)lower (rs −0.104 to −0.211)better
depression symptoms (MDI)lower (rs −0.105 to −0.149)better

So the same disposition is unpleasant to inhabit and associated with less trait psychopathology. That is not an inverted U, and it is not a taxonomy problem — it is one construct, measured once, yielding opposite verdicts against two different outcome classes.

What it does to resolution 1. The inverted-U reading needs interoceptive contact on a single outcome axis with an optimum somewhere along it. Here the axis forks: momentary valence and trait symptom burden are both plausible readings of “better off,” and body-wandering scores badly on the first and well on the second. Before asking where the optimum is, this page now has to ask optimal for what, over what timescale. That question was always implicit — Farb’s warning that awareness training can worsen anorexia initially is the same structure — but this is the first row to make it a measurement rather than a clinical caveat.

What it does to resolution 2. It helps, but less than the other new rows have. Resolution 2 sorts the collision by which construct is measured. Body-wandering is cleanly sensibility — self-reported attention, no detection, no ground truth — so it sorts neatly. But sorting it does not dissolve its internal split, because both of its verdicts come from the same construct. Resolution 2 explains disagreements between rows; this disagreement is inside one.

The reading the authors offer, and the wiki’s reason to take it seriously. Their proposal is that spontaneous bodily attention indexes preserved engagement with ongoing sensory input, protective relative to thought decoupled from the present body — and that negative affect at rest is therefore not equivalent to maladaptive cognition.

This wiki has met that claim before, from a different literature, in a different population, with a prospective design. Farb et al. (2011): elaborative self-referential reactivity predicts relapse, sensory reactivity protects, and felt sadness predicts nothing. The Farb row’s own stance line on this page reads “not how much, but which mode.” Banellis et al. supply that same dissociation in an unselected resting sample, and with the sensory pole finally interoceptive rather than visual — which is the gap the Farb row has carried as its standing caveat since 2011.

Two independent literatures converging on mode over magnitude, and on the specific claim that momentary unpleasantness is the wrong outcome variable, is the strongest structural support this page has for reorganizing around something other than dose. Recorded as convergence, not proof: Banellis et al. is cross-sectional, small-effect, and correlational, and it did not measure relapse or any clinical outcome at all.

What would sharpen it. The discriminating study is close at hand and has not been run: body-wandering scores and an actual interoceptive performance measure in the same participants. Both instruments exist in the same group’s hands (HRDT, RRST), and both cohorts come from the same Visceral Mind Project. Whether the person who drifts toward their body is also the person who reads it accurately is, on the evidence of the companion paper, quite likely to be no — and if so, this row and the heartbeat-task rows above it are not measuring the same thing at all, and this page’s table has been stacking incommensurable evidence for its entire life.

The row that answers a different question, and probably the better one

Every row on this page is a verdict about a quantity of interoceptive contact. Harrison et al. (2021) is the first source here to ask about a level instead, and the answer reorganizes the page more usefully than another verdict would.

They measured, in the same 60 people: four affective and four interoceptive questionnaires; respiratory perceptual threshold, decision bias, metacognitive bias and metacognitive efficiency; and peak anterior insula activity for four model-derived quantities. Sixteen measures spanning the whole hierarchy of one channel. Then they asked which of them carries the variance that separates low- from moderate-anxiety participants.

Loadings on the first principal component, in order:

  1. depression, state anxiety, anxiety sensitivity, anxiety-disorder symptoms
  2. breathing catastrophizing, negative MAIA interoceptive awareness
  3. negative metacognitive bias, body perception, negative metacognitive efficiency
  4. perceptual threshold, decision bias
  5. anterior insula activity, last

What this does to resolution 2. Resolution 2 says the question is malformed because “interoceptive awareness” names several dissociable constructs. This is the first source on the page to put most of those constructs in one sample and rank them against an affective outcome — and the ranking is monotone in distance from the body. Beliefs about the body carry the relationship; confidence carries some; actual detection carries little; the interoceptive cortex carries least. Resolution 2 predicted the constructs would dissociate. It did not predict they would order themselves this way.

And it explains a pattern this page has been recording as a series of anomalies. The rows built on questionnaires (Oldroyd, Banellis, Volynets) produce clean, sizeable effects; the rows built on task performance (Dunn, Pollatos, Wiens, Werner, Nentjes) produce small, contested ones. This page has treated that as a collision of findings. It may instead be one finding: sensibility measures index the level where affect lives, and accuracy measures do not — which is also what Banellis et al. concluded from the other direction, and what metacognitive-efficiency says about bias versus efficiency.

The caveat the authors raise themselves, and it is the right one. Higher levels may simply be measured with less noise than psychophysical thresholds and brain signals, so the ordering could be detectability rather than strength. Nothing in the study separates those. But note the shape of the concession: even on the deflationary reading, the practical advice is the same — if you want to detect a relationship between affect and interoception, measure beliefs, and the field’s decades of investment in detection tasks buys the least.

A third candidate resolution, arriving from the timescale

The two resolutions below are about dose and about construct. Wallman-Jones et al. (2023) make a case for a third dimension that neither of them contains, and it is worth separating because two rows now show it.

Their within-person effect of physical activity on self-reported interoception is positive; their between-person effect of the same variable on the same measure is negative, and larger. Banellis et al. two rows up report the same architecture with different content — body-wandering unpleasant in the moment, protective as a disposition. Two studies, two constructs, two labs, one shape: the momentary and the dispositional versions of a quantity can point opposite ways, and every other row on this page reports only one of the two.

That is not resolution 1 (there is no curve; there are two slopes with different signs) and it is not resolution 2 (the construct is held fixed — same instrument, same participants). It is a third malformation: this page has been asking “is more better?” of quantities that do not have a single value per person. state-vs-trait-interoception gives the underlying finding — roughly half the variance in self-reported interoception is within-person — and the arithmetic consequence for the rest of the table is uncomfortable. Every row above measured its construct once. If the construct has an ICC near 0.5, the correlations those rows disagree about are attenuated by an amount none of them estimated, and the sharpest disagreement on this page (Pollatos r = 0.34 vs Dunn r = .08) is between two single-session estimates of a moving target.

What stops this settling anything. The ICC estimate is for sensibility, on a mindfulness questionnaire, in 70 students. No one has partitioned the variance of an accuracy measure across occasions, and until someone does, the claim that the accuracy rows are attenuated is an inference from an adjacent construct. Recorded as the page’s newest structural worry, not as a correction to its arithmetic.

A fourth malformation: the measure is a sum of two things with opposite signs

The three resolutions so far say the question is malformed by dose (resolution 1), by construct (resolution 2), and by timescale (the section above). Ventura-Bort et al. (2021) add a fourth, and it is the most mundane and possibly the most damaging, because it applies within a single construct that resolution 2 already sorted.

Resolution 2 tells a reader to check whether a row measured accuracy, sensibility or regulation, and to stop comparing across them. Do that here, and every questionnaire row on this page — Oldroyd, Banellis, Wallman-Jones, Harrison’s top-loading measures, and much of the Farb applied literature — sorts into the single bin marked sensibility. Ventura-Bort et al. put eight such questionnaires in one sample and found the bin contains two orthogonal factors:

Sensibility factor (believed accuracy)Monitoring factor (deployed attention)
well-being (W-BQ12 total)r = .58*r = .12
negative well-beingr = −.54*r = .07
MAIA Not-worryingr = .40*r = −.20*
granularity, negative emotionsβ = +0.27β = −0.31
granularity, positive emotionsnsβ = −0.21
felt emotional intensityβ = −0.20 (trend)β = −0.31

Three of those rows have opposite signs, and two of the differences are formally tested (Z = 3.94, Z = 4.94).

What that does to the table at the top of this page. Any row whose instrument mixes the two factors reports a weighted average of a positive and a negative association, with weights set by which questionnaire the authors happened to use. The MAIA is exactly such an instrument — its Attention regulation and Trusting subscales load on one factor, its Noticing and Emotional awareness on the other. So the rows built on MAIA composites are not measuring a dose of one thing, and their disagreements with each other may be disagreements about item mix.

It is not a new claim, it is a measured version of the page’s oldest one. Oldroyd et al. said in 2019 that the anxious profile scores high on noticing and low on not-worrying, and that a composite would score it as good interoception. This page has cited that as a profile observation ever since. Ventura-Bort et al. show it is a factor structure: the two poles are not a quirk of the attachment-anxious, they are orthogonal dimensions on which everybody has a position.

And it sharpens the Harrison ranking two sections up. Harrison et al. found sensibility carrying more of the anxiety variance than accuracy or insula activity. Read together, the two say: sensibility outranks accuracy, and within sensibility, the belief half is where the good outcomes live and the attention half is where the bad ones do. The field’s own instruments, in other words, contain the answer to this page’s question — they just report it as one number.

What stops it settling anything. It measured no accuracy, so it cannot speak to any heartbeat-task row. It measured no clinical outcome — “well-being” is a 12-item questionnaire with α = 0.50 in this sample. Its two-factor solution was forced. And the causal direction of its headline effect is undetermined by design: the authors themselves suggest that poor differentiation may cause more monitoring rather than the reverse, which would make Monitoring a symptom rather than a disposition, and change the row’s advice entirely.

A fifth malformation, and the page’s first randomized evidence: the quantity has no valence until a demand is specified

The four malformations above — dose, construct, timescale, item mix — are all measurement critiques. Eusebio et al. (2022) add one that is not, and it arrives with the design this page has never had: random assignment.

Every other row here observes a quantity of interoceptive contact and reads a verdict off an association. This trial assigned people to two 10-week trainings built to be identical except in interoceptive emphasis, and then asked what trait predicted benefit in each.

Interoception-Focused armMovement-Focused arm
did the arm raise MAIA more?no — both arms rose equally (p = .443)
did the arm reduce symptoms more?no — both arms fell equally (p = .311)
who gained sustained attention?gain scaled with MAIA (p = .007); no other questionnaire (ps > .15)only the lowest baseline MAIA scorers (p = .019); high scorers gained nothing

The first row is the page’s cleanest null and should be stated plainly. On the literal question — is more interoceptive awareness better, and does training it help — a randomized trial with an active comparator found that a curriculum designed to cultivate interoceptive awareness produced no more interoceptive awareness and no more symptom relief than one that deliberately avoided mentioning the body. Most of this page’s applied rows rest on the assumption that this comparison would come out the other way.

The third row is why the null is not the end of it. The same trait predicts benefit in one arm and predicts absence of benefit in the other. Not a different magnitude — a different sign, from the same measure, in the same trial, set entirely by what the participant was instructed to do.

That is a genuinely different malformation from the four above. Resolution 1 (inverted U) needs contact on an outcome axis with an optimum. Resolution 2 needs the construct pinned down; it is pinned here, and pinned identically in both arms. The timescale malformation needs two timescales; there is one. The item-mix malformation applies (this is a MAIA composite) but does not explain a sign reversal across experimental conditions — an unknown mixture of Sensibility and Monitoring cannot be protective under one instruction and costly under another without something outside the instrument doing the work.

The something outside is the demand. Interoceptive sensibility is a disposition to attend inward; whether that helps depends on whether the task in front of you rewards attending inward. This is the Dunn gain-term reading with the multiplicand supplied by the environment rather than by the body: Dunn showed accuracy multiplies whatever the body happens to be signalling, and this shows sensibility multiplies whatever the situation happens to be asking for. Both say the same structural thing — interoception is a coefficient, not a term — from opposite directions, and it is the second time this page’s central question has dissolved into a moderation rather than resolving into a dose.

What stops it settling anything. N = 34 analysed against a power analysis calling for 36 — the interaction that carries the argument is the least well-powered thing in an underpowered trial. It is one trial, unreplicated, from a group with a theoretical stake in the interoceptive arm and a declared commercial conflict on one co-author. The outcome is SART accuracy, not wellbeing, which is the same complaint this page files against the Dunn and Werner rows. The moderator is a total MAIA. And the paper’s own concluding sentence misdescribes its results, which is not a reason to discount the analyses but is a reason to read them rather than the abstract. Recorded as the page’s first randomized evidence and its first manipulated moderation — not as the answer.

The two candidate resolutions

1. It is an inverted U. There is an optimal band of interoceptive contact, and both tails are dysfunctional — hypervigilance above, numbness below. Attractive, and consistent with the table, but no source in this wiki tests it: nobody has measured interoceptive contact and wellbeing across a full range in one sample and looked for curvature. It is currently a shape imposed on a scatter of studies, not a finding.

2. The question is malformed, because “interoceptive awareness” is not one thing. This is the stronger reply, and it is Farb et al.’s whole point. Sort the rows above by which construct is actually being measured and the contradiction thins:

  • Volynets measures the intensity of a state report — how much body gets coloured. Less signal.
  • Oldroyd measures sensibility via maia — beliefs about one’s interoceptive style. More attention, worse relationship.
  • Farb’s training targets regulation — the relationship to the signal, “without necessarily changing it.”
  • The heartbeat-detection-task literature measures accuracy, which meditators famously do not improve (Khalsa et al. 2008; Parkin et al. 2013) despite improving on everything else.

Under this reading, “less signal” (Volynets) and “better relationship to signal” (Farb) are not competing answers to one question; they are answers to different questions that share a word. An older adult with a quieter body and a good relationship to it, and an anxious young adult with a loud body and a frightened relationship to it, can both be doing well or badly depending on the term you pick — and the interoceptive-taxonomy predicts exactly this dissociability.

What would actually settle it

The reason to keep this open rather than dissolve it into (2): resolution 2 explains away the appearance of disagreement but leaves the clinical question standing. Farb et al. and Lyons et al. both recommend body-awareness training for depressed patients; Farb et al. also warn it can be harmful in severe and suicidal depression, whom Lyons et al. excluded. Two papers recommending the same intervention for adjacent populations, one of which says the intervention is dangerous for the population the other did not study, is a live disagreement no amount of taxonomy tidies away.

Discriminating tests the wiki does not have:

  • A dose-response or curvature analysis of interoceptive contact against wellbeing within a single sample — the direct test of resolution 1. Still absent, but narrowed: Dunn et al. (2010) ran the equivalent analysis for decision quality and found neither a curve nor a line (r = .08), only a moderation. Whoever runs the wellbeing version should look for the same shape — an interaction with whatever the body is actually signalling — rather than for curvature in a main effect that may not exist.
  • Any study that moves one taxonomy construct while holding the others fixed — the direct test of resolution 2.
  • An adjudication of the Pollatos/Dunn arousal disagreement, which is now the page’s most tractable open question because the two designs differ in only a few identifiable ways. The decisive study measures interoceptive accuracy, the bodily response, and felt arousal in one unselected sample large enough to estimate both the main effect and the interaction — which is Dunn’s design with Pollatos’s sample size problem removed. Neither existing study can do it: Pollatos et al. have no bodily-response term, and Dunn et al. have a sample in which the main effect, if it is the size Pollatos’s selected sample suggests, may simply be undetectable.
  • A replication of the Dunn moderation, and its wellbeing analogue. The gain-term reading now carries a lot of this page’s weight and rests on two unreplicated interactions in one 2010 paper about card games. If interoceptive accuracy multiplies bodily signal in decision-making, the obvious next question is whether it multiplies bodily signal in distress — which would predict that awareness training helps people whose bodies are giving good information and harms people whose bodies are lying, and that is testable in exactly the populations Farb et al. warn about (anorexia, panic, chronic pain), where the body is arguably wrong by definition.
  • Whether the ageing decline Volynets et al. observe is the same quantity that body-awareness training raises. If ageing reduces signal while training changes relationship-to-signal, they never touched.
  • Whether blunting is iatrogenic numbness or successful downregulation. Patients call it a side effect; a strict “less is more” position predicts they should feel better for it.

Caveats on the evidence, recorded so this page does not overclaim

The Lyons row is the weakest of the six. Its “less activation in depression” claim rests almost entirely on a single fear cell in a table whose ratio columns are transposed, and disappears when fear is excluded — see the arithmetic note on lyons-2021-body-maps-depression. What survives from that paper is the recommendation (body awareness as a treatment target) and the medication finding, not a clean demonstration that depression blunts the felt body.

The Volynets row is a proposal, not a result: the age effect is real (rs = 0.11, cross-sectional) but the chain from it to “easier to regulate” to “increased life satisfaction of old age” is speculation the authors offer in a discussion section, resting on Scheibe & Carstensen (2010) and Urry & Gross (2010) rather than on their own data.

So this debate is at present a collision of two well-measured findings and two under-evidenced interpretations of them. That is worth a page — the interpretations are what the applied literature runs on.