Desmedt et al. (2022) — HCT performance and mental-health outcomes: a meta-analysis

The empirical companion to Desmedt et al.’s (2023) conceptual critique, and the paper the wiki’s olivier-desmedt page had been flagging as “still queued in raw/.” Where the 2023 review argues a priori that the field mislabels its measures, this 2022 meta-analysis supplies the criterion-validity datum: the heartbeat-counting task — the instrument under nearly every “interoceptive accuracy” claim in this wiki — does not predict the mental-health outcomes the construct was built to explain, in 133 studies and 11,524 adults.

The design in one line

A pre-registered (PROSPERO CRD42019142176), PRISMA-compliant, random-effects meta-analysis of zero-order associations between HCT performance and DSM-relevant risk factors/indicators of mental health, restricted to the seven outcomes studied in at least ten published papers, with outlier detection, three-level models to handle within-study dependence, moderator tests (population type, measurement type), and publication-bias correction (funnel/Egger, p-curve, trim-and-fill). Data and R scripts are open on OSF.

The result that matters: a clean dissociation

outcomeN (studies)rsignificant?shared variance
trait anxiety400.03no~0.1%
depression31−0.04no (−0.08, p=.09 sans 1 outlier)0.6%
alexithymia23−0.01no0.01%
heart rate40−0.17yes2.9%
BMI29−0.11yes1.2%
sex (male>female)14−0.14yes2.0%
age20−0.06 → −0.07/−0.11only after correction<1%

The three psychological constructs — the ones the interoception→emotion→psychopathology story predicts should correlate with cardiac accuracy — are flat. The three that reach significance are properties of the cardiac signal or the body producing it: a faster heart, a higher BMI, and (on the field’s reading) male physiology all make the beat harder or easier to detect. This is exactly the cardiodynamic confound read at meta-analytic scale: the score tracks how loud the heart is, not how well the person perceives it or how they are doing mentally. The authors are careful that the biological associations cut two ways — a signal the perceiver genuinely detects (pro-validity) or a signal-intensity artefact (anti-validity) — and note that the HCT↔heart-rate correlation itself goes non-significant once time-estimation and heart-rate knowledge are controlled (Desmedt et al. 2020), which is the guessing account.

Why this is a criterion-validity blow, stated carefully

The paper is scrupulous that a null does not by itself condemn the task: the absence of an HCT↔mental-health association is jointly consistent with (a) no true relationship between cardiac interoceptive accuracy and these outcomes, (b) the HCT failing to measure cardiac IAcc, or (c) both, and a meta-analysis of one task cannot separate them. What it can do is remove the empirical motivation the task rode in on. The HCT became ubiquitous because interoceptive accuracy was theorised to matter for anxiety, depression and alexithymia (Craig 2004; Paulus & Stein 2006; Barrett et al. 2016) — and here, at the largest scale available, it does not track any of them. Combined with the direct experimental evidence the discussion assembles (feedback about one’s heart rate changes performance; pacemaker manipulations do not; modified “count only felt beats” instructions halve scores; response bias inflates them), the authors read the nulls as adding to the case that the task’s validity is the problem, and prescribe restricting conclusions to “the capacity to estimate heart rate via mental counting” rather than “interoceptive accuracy.”

Where it lands on the wiki

  • is-the-heartbeat-counting-task-valid. This is the strongest single addition to that debate since the Van der Does pooling: it converts the sceptical case from “the score is confounded” (a construct argument) into “the score predicts none of the things it was supposed to” (a criterion argument), and does so from inside the cardiac domain. It strengthens the Desmedt & Corneille field-position and supplies the outcome side that the guessing-mechanism papers (2018, 2020) did not.
  • alexithymia. A direct hit on the objective-task version of the Brewer/Bird “alexithymia is a general deficit of interoception” claim: at the level of HCT performance (not self-report), alexithymia and cardiac accuracy share 0.01% of variance across 23 studies. The wiki already held that the alexithymia↔interoception link survives mainly “at the level of beliefs” (Ventura-Bort et al.); this is the complementary objective null.
  • is-more-interoceptive-awareness-better. Another row where “more accuracy” fails to buy a measurable outcome — here, the negligible effect sizes explicitly undercut interventions that target interoceptive ability to reduce depression.
  • Aging. The age association (significant only after bias correction / in the three-level model) is the meta-analytic backdrop to Khalsa et al.’s first-hand discrimination-task age decline and MacCormack’s felt-body decline — consistent in direction, tiny in the counting task, and (per Murphy et al. 2018) partly mediated by BMI.

Caveats the wiki should keep attached

The authors themselves flag a positivity bias in their inclusion rule: by taking only the ≥10-study outcomes, they analysed the associations researchers thought most worth studying — so if anything the design was stacked toward finding effects, and still found none for the psychological outcomes. Against that, heterogeneity was high (I² up to 80.8%), covariates could not be modelled, and the guessing-reduced modified instructions were too rare to test as a moderator — so a real HCT↔mental-health relationship that emerges only under clean instructions or with the right covariates is not excluded. The honest summary is the authors’ own: five non-mutually-exclusive explanations for the nulls, and this study cannot say which, but it “highlights that theories or measures should be advanced to allow for strong a priori tests.”