Murphy et al. (2020): Testing the independence of self-reported interoceptive accuracy and attention

Six studies developing and validating the Interoceptive Accuracy Scale, from the Bird/Catmur lab with Khalsa as co-author. The stated aim is modest — build a self-report measure of believed interoceptive accuracy, because one barely exists — but the paper is doing something more consequential, and the wiki should read it for that: it is the first study here to test the sensibility split against an objective task.

The model being tested: a 2×2, not a list

interoceptive-taxonomy tracks five partitions of interoception, all of them flat lists of constructs. Murphy, Catmur & Bird (2018, 2019) propose a different shape, and this paper is its empirical test:

what is measured: accuracywhat is measured: attention
how: performanceheartbeat counting / detectionexperience sampling of bodily attention
how: self-report beliefconfidence ratings, IAS, ICQBPQ, MAIA Noticing

Two crossed factors: what is assessed (can you perceive accurately vs is the body the object of attention) and how (objective performance vs belief). The point of the crossing is a prediction the flat lists cannot make. Measures should relate when they share the what, whichever the how — so a self-report of believed accuracy should track objective accuracy, and a self-report of attention should not.

The quadrant with almost nothing in it was self-reported accuracy: existing questionnaires either measure attention (BPQ, MAIA) or bundle the two together (MAIA, Body Awareness Questionnaire). The IAS was built to fill it, which is what makes the test possible at all.

What they found

Self-report splits exactly along the what axis. In three independent samples, the two accuracy-belief scales (IAS and the reverse-scored ICQ) correlate at r = .57–.71 while the attention scale (BPQ) correlates with neither. This is Ventura-Bort et al.’s (2021) believed-accuracy/deployed-attention factor split arriving from a different lab, a different country and a different method — correlation structure rather than forced PCA — with the same seam.

And the how axis is crossed successfully. Heartbeat counting accuracy correlated with the IAS (r = .271, .336) and with nothing else administered: not the BPQ, not the ICQ, not the TAS-20, not time-estimation performance. Post-task confidence — the state version of believed accuracy — behaved identically, correlating with the IAS and with HCT performance but not with the BPQ.

That is the 2×2’s central prediction confirmed on both axes, in the paper built to test it. Its value to the wiki is not the scale.

Why this matters more than the paper claims: the wiki’s most durable null gets a boundary

Across interoceptive-sensitivity, interoceptive-taxonomy, maia and heartbeat-detection-task the wiki records the same anomaly over and over: self-reported interoception does not correlate with objective interoceptive accuracy. Meditators attend without detecting. Attachment-anxious participants notice without perceiving. Harrison et al. find sensibility outranking accuracy against anxiety with a sparse cross-block correlation matrix. Banellis et al. find the generalizing component of confidence is not interoceptive at all. Wallman-Jones et al. produced the wiki’s first positive association and it was filed as friction.

This paper supplies the boundary condition those nulls were missing, and it is not subtle. Every one of those nulls used an attention-flavoured self-report — MAIA Noticing, the BPQ, generic body-awareness scales — against an accuracy-flavoured task. Match the what, and the correlation appears; mismatch it, and it does not, in the same participants, in the same session, twice.

Ventura-Bort et al.’s page closes by naming the discriminating study: “run these questionnaires alongside an objective interoceptive measure and a physiological one.” Half of that study is this one, and it was published first. The wiki’s sensibility↔accuracy dissociation is therefore, on current evidence, an attention↔accuracy dissociation — which is not the same claim, and is far less damaging to the taxonomies than the version the wiki has been carrying.

Two brakes worth applying before that reframing hardens.

The effect is small and the samples are tiny. r = .27 and r = .34 at n = 52 and n = 33, p = .047 and p = .049. Roughly 7–11% shared variance, in studies with the statistical power to detect little else. Nothing here rules out the reframing being partly a chance separation of two weak correlations, one of which happened to clear .05.

The objective anchor is the contested one. If the HCT is largely guessing from a believed heart rate, then a correlation between the HCT and a questionnaire asking “can you accurately perceive your heartbeat?” is not obviously a perception result. The paper’s defence against this is the control battery, and see below for why that defence is weaker in this dataset than in the literature it cites.

What it contributes to the heartbeat-task debate — honestly scored

This is the wiki’s first pro-validity empirical datum on the HCT in some time, and it should be recorded as one without being oversold.

For validity. Under a pure guessing account, HCT performance is driven by beliefs about resting heart rate plus time-estimation skill. Neither of those is a plausible cause of the selective pattern found here: time-estimation task performance correlated with no questionnaire, the HCT correlated with the accuracy self-report and not the attention self-report, and the relationship held with estimated resting heart rate partialled out. A guessing account has to explain why a belief-driven score tracks one belief questionnaire and not the other.

Against. The paper’s own control-variable analyses failed to replicate: in Study 5, HCT accuracy was unrelated to heart-rate beliefs and to time estimation, contrary to Ring et al. (2015) and to Murphy, Brewer et al. (2018) — the authors’ own paper, the one that established the control battery. Controls that show no effect cannot demonstrate that a relationship survives them. And the enormous HCT-confidence↔HCT-accuracy correlation in Study 5 (r = .806) is exactly as consistent with both quantities flowing from one shared belief about one’s own heart as with genuine metacognitive insight.

Net: a real point for the defenders, at n = 35, with its own controls non-replicating. See is-the-heartbeat-counting-task-valid.

Alexithymia lands on the accuracy side

alexithymia on this wiki has an awkward split: the TAS-20 factors with bodily belief scales (Ventura-Bort) while the objective-task version of the claim is dead flat (Desmedt et al. 2022, r = −.01). This paper sharpens the belief half into something more specific than shared method variance: the TAS-20 tracks the accuracy self-reports (r = −.43, −.57 with the IAS; .65 with the ICQ) and not the attention self-report (r ≈ .07 with the BPQ), holds after partialling self-esteem, and holds independently of depression and anxiety in regression — while the BPQ, in the same regressions, is predicted by anxiety and nothing else.

So “alexithymia is a general deficit of interoception” is, at the level of beliefs, a claim about believed accuracy, not about how much attention goes to the body. Which is a cleaner claim, and still a claim about beliefs: the objective-task null stands untouched, and this paper’s own HCT data show no TAS-20↔HCT relationship either (all ps > .13), consistent with Desmedt.

What the BPQ turns out to be

A byproduct worth keeping, because the BPQ is used across this wiki as a general interoceptive-sensibility scale. Here it is: excellent internal consistency (α = .98), good test–retest (r = .684, though poorer than the ICQ’s .814, Z = 2.282, p = .022), an imperfect one-factor CFA fit (RMSEA = .085, CFI = .877), uncorrelated with objective cardiac accuracy, uncorrelated with the two accuracy self-reports, uncorrelated with alexithymia — and predicted by anxiety alone, an effect that did not survive robust regression. Reliable, and measuring something narrow: how much of the body is attended, with a distress component.

Placement in the wiki

Creates interoceptive-accuracy-scale as a method page — the instrument Ventura-Bort et al. used and the wiki had named without hosting. It is the second Murphy source in raw/ and the first with data: her 2022 commentary argues that ability is not use, and this paper is the measurement work that argument stands on. It refines rather than refutes the accuracy/sensibility dissociation on interoceptive-sensitivity and interoceptive-taxonomy, adds a defender’s datum to is-the-heartbeat-counting-task-valid, and narrows alexithymia’s belief-level claim. Held at its actual weight: a careful psychometric paper whose most important result rests on two correlations of r ≈ .3 in samples of 35 and 52.