Is the heartbeat-counting task a valid measure of cardiac perception?
The wiki’s most load-bearing methodological question, created with the Van der Does et al. (2000) ingest — and overdue, because heartbeat-detection-task had been accumulating objections in a limitations list for eight ingests without anyone naming the fight they belong to.
Why it matters more than its host literature suggests. This debate lives in 1990s panic research and is conducted almost entirely in terms of panic disorder. But Seth, Dunn, Pollatos, Oldroyd and the interoceptive-taxonomy’s entire “accuracy” construct rest on this instrument. If the sceptical position is right, a large fraction of this wiki’s quantitative evidence is a correlation computed across a mixture in which most participants contribute noise.
What is agreed
More than most debates here, this one has a shared dataset. Both camps contributed their raw data to the 2000 reanalysis and both are authors on it. Nobody disputes:
- Accurate perception is uncommon. 17.1% of panic patients, 7.9% of normal controls, 0% of depressed patients meet a <10%-error criterion. In every group, most people cannot do the task.
- Nearly everyone produces a number anyway. Over 95% report perceiving their heart rate; only 3.5% report feeling nothing.
- Almost everyone undercounts. Independently corroborated by Zoellner & Craske (1999): 89.7% of infrequent panickers and 88.9% of controls underestimated, ~7–8% overcounted — the same near-universal undercounting, from an outside lab. (They also independently replicated the panic-group accuracy advantage, which is a point for Ehlers on the group difference; see below.)
- Stroke volume predicts performance (Schandry et al. 1993).
- Raising heart rate by exercise raises measured accuracy transiently — above ~100 bpm, decaying to baseline by ~95 bpm, equally in patients and controls.
The disagreement is entirely about what those facts mean.
A second task family agrees on the prevalence
The debate above is conducted entirely on the counting task, and a natural escape for the validity defender is that the rarity of accurate perception is a counting artefact — solved by switching to a discrimination method that cannot be faked by guessing from a believed heart rate. Wiens, Mezzacappa & Katkin (2000) closes that escape. Their preferred-interval discrimination task — tones at fixed 200 ms or 500 ms delays after the R-wave, forced-choice — classified only 9 of 52 unselected undergraduates (17.3%) as accurate detectors, essentially the counting prevalence in the pool above (7.9% controls, 17.1% panic). So “most people cannot do it” is not specific to counting: a task designed to remove counting’s central confound finds the same rare accurate minority. That is a point for Van der Does on the prevalence fact, and against reading the rarity as a scoring artefact.
But it cuts both ways, and the wiki should record both edges. In that accurate minority, cardiac discrimination did track something real — good detectors felt emotion more intensely, with sympathetic arousal controlled (see wiens-2000-heartbeat-detection-emotion, core-affect). So the discrimination task supports the sceptics on how many people perceive accurately while supporting the defenders on whether accurate perception, where it exists, is a real ability that does psychological work. The two task families converge on a threshold picture: genuine cardiac perception is rare, and where present, it matters. What neither escapes is the cardiodynamic confound — a louder heart is easier to time as well as to count, and Wiens et al. measured skin conductance, not stroke volume.
The crux: what does undercounting show?
Both sides build on the same observation and reach opposite conclusions, which is what makes this a real debate rather than a disagreement about evidence.
| Ehlers | Van der Does | |
|---|---|---|
| Why do people undercount? | They perceive their heartbeats and miss a few. Logical, expected, evidence of validity. | They feel a regular rhythm slower than their actual HR — which is what participants actually report. Not missed beats: a different rhythm. |
| What does stroke volume predicting performance show? | The task tracks real cardiac events — evidence for validity. | (Not addressed directly in 2000; this wiki reads it as evidence the score indexes signal amplitude, not skill.) |
| What is the right score? | Continuous % error. A dichotomy imposes an artificial boundary on a continuum. | Categorical. The boundary is real, because it separates people the task measures from people it does not. |
| What is the artefact risk? | Time estimation — and it has been ruled out. | Anxious patients expect a faster rhythm, count faster, and land closer to truth because everyone undercounts. Lower error through anxiety, not perception. |
The self-report evidence is the sharpest thing either side has, and it favours Van der Does. “I lost count of a few” and “I felt a steady rhythm that turned out to be slower than my heart” are different experiences, and participants report the second. Ehlers’s validity argument requires the first.
An anxiety manipulation that predicts the artefact — from a lab that read it the other way
The crux above is usually argued on the undercounting fact alone. Zoellner & Craske (1999) add a within-subject test of the artefact prediction, and it is the sharper form of the evidence because it is a manipulation rather than a correlation.
They stratified each participant’s own trials into their lowest, medium and highest state anxiety and scored heartbeat error at each. Error fell as anxiety rose — greatest at the lowest-anxiety trials (54.7%), least at the highest (45.8%), main effect F(2,81)=10.04, p<.001 — with no sign of the inverted-U detriment they had predicted at extreme anxiety. Anxiety only ever improved the score.
That is exactly what Van der Does’s artefact predicts and nothing else needs to be true for it. Error here is unsigned (|actual − perceived|/actual), and in a task where ~89% of Zoellner & Craske’s own participants undercounted, any process that makes someone count more moves them toward truth and lowers error with no improvement in perception. An anxious person who expects a racing heart counts faster and scores as more accurate — the artefact is a within-subject anxiety effect, and here is one.
The instructive part is that Zoellner & Craske read their own result the opposite way — as state anxiety magnifying attention to bodily cues (the attentional-bias tradition), i.e. anxiety sharpening genuine perception. Both readings predict the identical finding, and the study cannot separate them, because it reports unsigned error and never checks whether the anxious trials’ counts shifted up toward truth (artefact) or simply became more accurate (perception). So the pro-validity lineage supplied, without intending to, a clean demonstration of the mechanism the sceptics propose — and interpreted it as perception. Recorded as a case where the two measurement models are observationally indistinguishable on the same data; the signed-error analysis that would separate them has not been run.
A stressor that raises the count while raising the heart
Schulz et al. (2013) add the manipulation version of the debate’s central worry: run one acute cold-pressor stressor through counting and discrimination in the same 42 people, and the counting score rises while the visual discrimination score falls (time×stress F[1,40]=4.17, p<.05; three-way F[1,40]=5.85, p=.02). Whether “acute stress improves cardiac interoception” is true depends entirely on which task you pick — which is the whole debate, demonstrated experimentally rather than argued from undercounting.
It bears on the crux the same way the exercise result does, and from the sceptics’ side. The cold pressor raised heart rate +8.2 bpm and systolic pressure +20.6 mmHg — sympathetic drive that increases contractility and makes the beat a louder signal — and the counting gain tracks it. So this is a second manipulation (after Antony et al.’s exercise) in which making the heart louder raises the score, with no perceptual learning available to explain it, and the authors concede as much because they did not measure stroke volume. Ehlers’s use of the stroke-volume correlation as evidence of validity meets the same rejoinder here: the task tracked the heart, not the perceiver. What Schulz adds beyond exercise is that the same sympathetic shift hurt a discrimination task keyed to a resting cardiac cycle (stress-shortened pre-ejection period desynchronizes the fixed R-wave-to-tone delay), so one physiological change pushes the two paradigms apart — the cleanest single reason on the wiki not to treat “counting” and “discrimination” as interchangeable operationalizations of one ability.
The criterion-validity turn: the score predicts none of what it was built to predict
Everything above argues about construct validity — what the score is, inferred from undercounting, self-reports, and manipulations of the signal. Desmedt et al. (2022) open a second, largely independent front: criterion validity — what the score predicts. If HCT performance indexes cardiac interoceptive accuracy, and interoceptive accuracy is theorised to matter for anxiety, depression and alexithymia (which is why the task became ubiquitous), then those associations should appear. Across 133 studies and 11,524 adults, they do not: trait anxiety r = 0.03, depression r = −0.04, alexithymia r = −0.01, all non-significant, none with detectable publication bias.
What reaches significance is diagnostic. HCT performance tracks heart rate (r = −0.17), BMI (r = −0.11) and sex (male > female, r = −0.14) — a faster, fattier-insulated, or physiologically male heart. These are the cardiodynamic confound and the signal-intensity account written across the whole literature: the score follows properties of the signal, not indicators of the mind the construct was supposed to bear on. And the HCT↔heart-rate correlation itself dissolves once time-estimation and heart-rate knowledge are partialled out (Desmedt et al. 2020), which routes even the “significant” associations back through guessing.
This is why the meta-analysis is a genuine addition rather than a restatement. The construct arguments on this page can each be met with “the score is noisy but still a valid continuum” (Ehlers’s move). A criterion null is harder to absorb: a valid-but-noisy measure of a real ability should still show attenuated true associations, not flat ones, in a sample this size. Desmedt et al. are careful that the nulls are jointly consistent with (a) no true IAcc↔mental-health link, (b) the HCT not measuring IAcc, or (c) both — a single meta-analysis of one task cannot separate them. But they remove the empirical motive the task was adopted for, and the same authors’ experimental guessing evidence tips the reading toward (b). The honest scorecard: this strengthens the sceptical side on prediction the way Van der Does strengthened it on prevalence — from within the cardiac domain, at the largest scale available, with open data.
The one caveat that keeps it from being decisive is the authors’ own: their ≥10-study inclusion rule biases toward the most-studied (hence most-expected-to-be-real) associations, so the nulls are not from a low-power fishing expedition — but heterogeneity was high, covariates were unmodelled, and the guessing-reduced modified HCT instructions were too rare to test, so a relationship visible only under clean instructions is not excluded.
The exercise result is the strongest evidence in the debate, and it arrived by accident
Antony et al. (1995) exercised participants and re-measured across seven trials while HR decayed from ~130 to ~92 bpm. Accuracy tracked the heart rate and nothing else: it rose while the heart was loud, returned to baseline by ~95 bpm, and was identical across PD, social phobia and controls. Of 60 participants, 25 showed the transient gain and one became durably accurate.
This converts the cardiodynamic confound from a correlation into a manipulation. Make the signal bigger, and accurate perceivers appear; let it fade, and they dissolve. No learning, no skill, no group specificity.
It also cuts against Ehlers’s own use of the stroke-volume finding, which is the neat part. She cites the cardiodynamic correlation as evidence the task tracks real cardiac events — and it does. But “tracks real cardiac events” and “measures the participant’s perceptual ability” come apart exactly here: a task whose score you can raise by making the heart beat harder is tracking the heart, not the perceiver. The validity argument survives; the construct it was defending does not.
And it embarrasses the sceptics slightly too. If the count were pure schema with no cardiac input, why would it improve when the heart gets louder? The exercise result requires that real cardiac signal reaches perception at least sometimes, for at least some people, at sufficient amplitude. The honest reading is a threshold: the heart is perceptible when loud enough, and most laboratory hearts are not loud enough, so most laboratory counts are something else.
What is unresolved, and what would resolve it
The categorical claim has never been tested. Van der Does et al.’s whole enterprise needs the error distribution to be a mixture rather than a continuum, and the evidence offered is three histograms and an eyeball. Fig. 1 is described as “normally distributed, with a marked and skewed peak around 0% error” — which is a description of a continuum with a spike, not of two populations. No dip test, no mixture model, no formal test of unimodality. This is a one-analysis question and the analysis has not been run, at least not in anything this wiki has read.
Discriminating tests the wiki does not have:
- A formal mixture/unimodality analysis of heartbeat-counting error scores in a large unselected sample. If the distribution is a mixture, the number of components and the boundary are estimable rather than stipulated, and the 10% convention can be replaced with a posterior probability. This is the single highest-value missing analysis on this page.
- Does the “inaccurate” count correlate with anything cardiac at all? Van der Does et al. (1997) reported that in inaccurate perceivers, perceived HR was unrelated to actual HR — which is the sceptical position’s core empirical claim and, in the 2000 pool, actual HR does correlate with counted beats at around 0.30 in every group (which is why it is partialled out of Table 4). Those two facts sit uneasily together and the paper does not reconcile them.
- Whether the minority-validity thesis holds outside cardiac interoception. Respiratory and gastric measures are untouched here.
- Whether “accurate perceiver” is a stable enough category to select on. The treatment data say less than half of accurate PD patients stay accurate across sessions. Extreme-groups designs like Pollatos et al.’s select on one session and analyse as though selecting on a trait.
The consequence the rest of the wiki has to absorb
If the sceptical reading is even partly right, the wiki’s live Dunn/Pollatos disagreement gets a new explanation that the wiki did not have and that points opposite to its current one.
The wiki’s account has been sampling: Pollatos selected 22 good perceivers from ~140, and extreme groups inflate correlations, so their r = 0.34 is not comparable to Dunn’s r = .08.
The mixture reading says the same fact runs the other way too: selecting the top ~16% is approximately selecting the subpopulation the task is valid for, and Dunn’s unselected sample dilutes the moderator with ~80% noise, attenuating any real relationship toward zero. Selection as purification, not just distortion.
Both can be true at once — selection can isolate a real subgroup and inflate the correlation within it — and the wiki now records both. See is-more-interoceptive-awareness-better, interoceptive-sensitivity, heartbeat-detection-task, where the disagreement is tracked.
The successor instrument exists, and it settles less than expected
This debate is conducted entirely on 1980s instruments. The HRDT (Legrand et al. 2022) is what the field built in response to precisely the complaint above: it replaces the count with a faster/slower judgment against a tone train in known units, and fits a psychometric function, so sensitivity and response bias are estimated separately rather than confounded in one error score. That is the signal-detection objection answered directly, and it also delivers a signed threshold — so the systematic under-estimation this whole debate turns on becomes a measured parameter with a direction instead of an inference from unsigned error.
Three things it does not do, and they map onto the unresolved list above.
It does not touch the cardiodynamic confound. A louder heart is easier to judge as well as easier to count. Everything in the exercise result and the stroke-volume finding applies unchanged, because those are facts about the signal, not the response format. The successor fixes the half of the problem that was about the participant’s answer and none of the half that was about their heart.
It does not test the mixture claim — but it finally could. The single highest-value missing analysis named above is a formal mixture/unimodality test of the error distribution in a large unselected sample. HRDT thresholds are continuous, in physical units, and Banellis et al. (2026) have released N = 513 of them publicly. The analysis Van der Does’s position has waited a quarter-century for is now a day’s work on an open dataset, and nobody in the wiki’s literature has run it.
And it inherits a new problem at the metacognitive layer. Cardiac M-Ratio was unestimable for 96 of 241 participants in that study, against 10 for the respiratory task. M-Ratio degrades when first-order performance approaches chance — and near-chance is where cardiac interoception lives, which is this debate’s founding observation restated in a modern statistic. See metacognitive-efficiency.
The larger relocation. The debate above asks whether the cardiac instrument works. is-interoception-domain-general, created with the Banellis ingest, asks whether the construct travels between organs, and the answer there is currently no. That is the more damaging of the two questions for this wiki, and it is independent of this one: a perfectly valid cardiac instrument measuring a real cardiac ability would still not license any of the general conclusions drawn from it.
A convergent-validity test the guessing account has to answer
Every argument above is about the score’s construct validity (what it is) or its criterion validity (what it predicts). Murphy et al. (2020) add a third kind, and it is the first evidence in some time that runs the defenders’ way.
The logic is a 2×2. Cross what a measure assesses (accuracy of perception vs the body as an object of attention) with how (objective performance vs self-reported belief), and a specific pattern follows: measures sharing the what should converge across the how. Having built the missing cell’s instrument — the IAS, a trait self-report of believed accuracy — they tested it.
| measured alongside HCT accuracy | correlation |
|---|---|
| IAS (self-reported accuracy) | r = .271 (n = 52, p = .047); r = .336 (n = 33, p = .049) |
| post-task confidence (state self-reported accuracy) | r = .507, .454, .806 |
| BPQ (self-reported attention) | ns, all ps > .13 |
| ICQ (self-reported accuracy, other instrument) | ns |
| TAS-20 (alexithymia) | ns |
| time-estimation task performance | ns with every questionnaire |
Why this is awkward for a pure guessing account. If the count is a belief about one’s own heart rate plus time-estimation skill, the natural prediction is that it tracks neither body questionnaire (the beliefs are about rate, not about perceptual skill) or, at a stretch, both. It tracked one and not the other, selectively, twice — and time-estimation performance, the guessing account’s other engine, correlated with nothing at all. The relationship also held with estimated resting heart rate partialled out (Study 5 supplementary), which is the sharpest form of the control.
Three reasons it settles less than it looks like it should.
The controls did not work in this dataset. HCT accuracy was unrelated to heart-rate beliefs and to time estimation in Study 5 — contrary to Ring et al. (2015) and to Murphy, Brewer et al. (2018), the authors’ own paper and the source of the control battery. A relationship that “survives” controls with no effect has not survived much. The authors flag the non-replication themselves and attribute it to sample size and characteristics.
The confidence–accuracy correlation is too good. r = .806 between mean confidence and counting accuracy in Study 5 is a striking number, and it is exactly what you would expect if both quantities are downstream of one shared belief about one’s own heart. The metacognitive-insight reading and the shared-prior reading make identical predictions here, which is the same observational-indistinguishability problem the Zoellner & Craske section above records.
The size. r ≈ .27–.34 at n = 33 and n = 52, both at p ≈ .047. About 7–11% shared variance. Against Desmedt et al.’s 11,524 participants on the criterion side, this is a small stone on the other pan.
Where it leaves the scorecard. The sceptics hold prevalence (Van der Does, Wiens), manipulability (Schulz, Antony), and criterion validity at scale (Desmedt). The defenders now hold one selective convergent-validity result, which is a kind of evidence the sceptical case had not previously had to absorb — and which the sceptics can absorb, if the IAS is measuring beliefs about the body that happen to be shaped by the same cardiodynamic facts that shape the count. A person whose heart is a loud signal both counts better and has learned they can tell when it is racing. That account predicts everything in the table, including the selectivity, with no perception in it — and it is testable, since it predicts the IAS↔HCT relationship should weaken when stroke volume is controlled. Nobody has measured stroke volume in any study on this page.
Why this is filed as open rather than as a limitation
Because the sceptical position is not a caveat, it is a rival measurement model with different implications for analysis (categorical vs continuous), sampling (select vs recruit), and inference (a mixture cannot be correlated across). And because the field went the other way: post-2000 interoception research overwhelmingly reports continuous heartbeat-counting scores from unselected samples, which is Ehlers’s practice, without engaging the argument that the practice is unsound. The debate did not resolve. It was left in the panic literature while the instrument moved into interoception research without it.