Heartbeat detection task
The canonical measure of interoceptive-sensitivity, originating with Schandry (1981), “Heart beat perception and emotional experience.” Participants report perceived heartbeats over intervals; the score indexes cardiac interoceptive accuracy.
The validity problem, which belongs at the top of this page
Everything below this section describes what the task predicts. This section is about whether the task measures anything, in most of the people it is given to — and it arrives late, with the Van der Does et al. (2000) ingest, because the page had been collecting objections in a list without naming the dispute they belong to. That dispute now has a page: is-the-heartbeat-counting-task-valid.
The prevalence datum. Seven pooled studies, 709 participants, re-scored uniformly:
| group | n | accurate (<10% error) |
|---|---|---|
| panic disorder | 275 | 17.1% |
| normal controls | 191 | 7.9% |
| palpitation patients | 99 | 6.1% |
| mood disorder | 32 | 0.0% |
Meanwhile more than 95% of participants produce a count, and only 3.5% report feeling nothing. So in the healthy adult samples that generate nearly all of this wiki’s interoception evidence, roughly nine in ten participants supply a number that is wrong by about 30% — and the interesting question is what that number is.
It is probably not a noisy estimate of heart rate. Participants do not report losing count. They report feeling a regular rhythm somewhat slower than their actual heart rate — a confident, steady, wrong percept, which is not what missing beats feels like. Van der Does et al. propose it is schema-guided: a reading of expectation rather than of the heart.
Why this is not just a reliability complaint. If the inaccurate majority are estimating a different quantity rather than estimating heart rate badly, then the continuous score is a mixture of two populations, and a correlation computed across an unselected sample is not an attenuated estimate of a real relationship — it is a statistic with no clean interpretation. That changes what analyses are licensed, not just how wide the error bars are. Ehlers disputes exactly this and defends the continuous score; see is-the-heartbeat-counting-task-valid for both positions stated fairly.
And it has never been properly tested. The mixture claim needs the error distribution to be bimodal, and the evidence offered is three histograms and an eyeball — Fig. 1 is described as “normally distributed, with a marked and skewed peak around 0% error,” which describes a continuum with a spike. No dip test, no mixture model. The task’s most serious challenge is one analysis away from being made properly, and nothing this wiki has read has run it.
The task’s heaviest users now flag its generalizability too. Quadt, Critchley & Garfinkel (2018) — the Sussex group whose dimensional framework and clinical programme run on this task — write that most current work depends on heartbeat tracking, and that its validity as a proxy “needs to be treated with caution”: beliefs about heart rate influence it (Ring & Brener 1996), tracking and discrimination performance can diverge (Ring & Brener 2018 report “heartbeat counting is unrelated to heartbeat detection”), the two tasks tap different processes, and the relationship of cardiac perception to respiratory and gastric axes is “scarce and inconsistent,” so cardiac findings may not generalize across the body. They add an IPP reason the cardiac task may under-generalize that is about the object rather than the measure: modalities may be precision-weighted differently, and in conditions where cardiac sensations matter less than others (e.g. eating disorders) the heart is the wrong channel to probe. See is-the-heartbeat-counting-task-valid for this recorded as a field-position.
The cardiodynamic confound is no longer a correlation
The section further down this page carries Schandry et al.’s (1993) stroke-volume finding as the objection that “keeps regenerating.” Van der Does et al. (2000) upgrade it from a correlation to a manipulation, and this is the most consequential single result on this page.
Antony et al. (1995) exercised participants and re-ran the task over seven trials while heart rate decayed:
| trial | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|
| actual HR (bpm) | 130 | 104 | 95 | 93 | 90 | 91 | 92 |
Measured accuracy rose while the heart was loud and returned to baseline by ~95 bpm. 25 of 60 participants showed the transient gain; exactly one became durably accurate. The effect was identical in panic patients, social phobics and controls, who did not differ in actual HR at any trial.
Nobody learned to perceive. The signal got louder and the score went up. Raise the amplitude of the thing to be detected and accurate perceivers appear; let it fall and they dissolve. This is what a signal-strength account predicts and what a perceptual-skill account does not.
Two riders. It corrects the original study’s own conclusion: Antony et al. had read the data as a general improvement in % error, but inaccurate perceivers’ post-exercise scores were no better than baseline — the mean moved because a subpopulation changed state, which is the categorical thesis demonstrated on the continuous score’s home ground. And it constrains the sceptics too: if the count were pure schema, why would it improve when the heart gets louder? The honest reading is a threshold. Real cardiac signal reaches perception when loud enough; most laboratory hearts are not loud enough.
The scores disagree with each other
The sharpest form the measurement worry takes, because it shows the choice of score changes the conclusion — on the same participants, in the same data:
| manipulation | continuous % error says | categorical scoring says |
|---|---|---|
| distraction (tone pips) | minimal effect (Ehlers et al. 1995) | 35% of accurate perceivers become inaccurate; 20% more drop to ‘probable’ |
| exercise | general improvement across participants (Antony et al. 1995) | no improvement in inaccurate perceivers; a subgroup transiently converts |
| treatment (pre/post) | stable individual characteristic (Ehlers & Breuer 1996; Antony et al. 1994) | fewer than half of accurate patients stay accurate |
Three manipulations, three reversals. Whatever else is true, “which score you use” is not a neutral analytic choice here.
The distraction result has a further consequence the authors flag: it likely explains why the rival method — matching one’s heartbeat to a series of tone pips (Brener & Kluvitse 1988; Whitehead) — “typically finds poor (chance level) performance and no differences among groups.” The tones are not a neutral response format. They are an interference that degrades the thing being measured. That is worth knowing before treating discrimination variants as a clean alternative to counting.
A stressor makes them disagree in opposite directions. Schulz et al. (2013) add the sharpest version of “which task you use changes the conclusion”: one acute manipulation (a cold pressor), two tasks, in the same 42 people, moving the score opposite ways. Counting accuracy rose after stress (time×stress F[1,40]=4.17, p<.05); visual discrimination accuracy fell (three-way interaction F[1,40]=5.85, p=.02). So the answer to “does acute stress improve cardiac interoception?” is instrument-dependent — there is no paradigm-free fact of the matter. And the study reads as a live demonstration of the cardiodynamic confound below rather than of attention: the cold pressor raised heart rate +8.2 bpm and systolic pressure +20.6 mmHg, so the counting gain is what a louder heart predicts, and the authors — unable to rule it out, having not measured stroke volume — say so themselves. In that same sample the two paradigms did not intercorrelate (Schandry vs Whitehead r=.22 and .08, ns) while the two Whitehead modalities did (r=.63), a first-hand instance of the Ring & Brener divergence Quadt et al. cite below.
The discrimination task: a different alternative, and what it independently shows
The tone-pip matching method above is not the only discrimination variant. The Katkin lab’s preferred-interval method — tones triggered by your own R-waves at a fixed 200 ms or 500 ms delay, forced-choice “simultaneous or delayed?” — is the family’s other branch, and it is the one built specifically to remove counting’s guessing confound: the tones carry your true rate and rhythm, so a belief about your resting heart rate buys you nothing, and only actual perception of beat timing lets you tell the intervals apart (Ring & Brener 1996). Wiens, Mezzacappa & Katkin (2000) is the wiki’s first first-hand use of it, and it lands on this page’s arguments in three places:
- Prevalence generalizes across task family. Only 17.3% of unselected undergraduates met the discrimination criterion — essentially the Van der Does counting prevalence (7.9% controls / 17.1% panic, <10% error). So “most people cannot do it” is not a counting artefact; genuine cardiac accuracy is rare on the discrimination task too. This strengthens the sceptical position on prevalence while showing it is not method-specific.
- The emotion effect survives arousal control — which no counting study here managed. Good detectors felt emotion more intensely across amusement, anger and fear (F(1, 50) = 7.61, p < .01) and no more pleasantly (F < 1), and the intensity effect held with skin conductance covaried out, because good and poor detectors did not differ in electrodermal activity or heart rate. This is the one study in the cluster that measured bodily arousal and showed the emotion relationship is not an arousal confound — the omission this page holds against Pollatos. See core-affect.
- But it does not escape the cardiodynamic confound. Wiens et al. controlled sympathetic arousal (SCR), not stroke volume. A more forceful heart is an easier signal to time as well as to count, so the signal-amplitude worry is orthogonal to counting-vs-discrimination and survives the switch. The discrimination task fixes the guessing confound and the arousal confound; it does not fix the loudness confound.
Do not read this against Ehlers (1993), who reports the panic-patient group difference is absent on Katkin discrimination. Wiens et al. measured a healthy-sample individual-differences correlation, not a group difference — different question, no contradiction, and arguably the informative pattern (the task tracks how intensely healthy people feel while not separating patients from controls, which is what an attentional/belief account of the panic effect would predict).
The Brener–Kluvitse simultaneity variant, read forensically
The discrimination family has a second branch this wiki now sees first-hand: the Brener & Kluvitse (1988) simultaneity-judgment method, used by Nentjes et al. (2013) in 75 male offenders. A tone is delivered either 250 ms (perceived simultaneous) or 650 ms (perceived nonsimultaneous) after each R-wave; the participant judges, forced-choice, whether the tones fell on their heartbeats; performance is scored with signal-detection d’ = z(HITS) − z(FA). It is the branch this page’s exercise section names as the one that “typically finds poor (chance level) performance,” and Nentjes et al. confirm the characterization at the sample level: mean d’ was 0.00 (SD 1.18) — chance. The forensic sample, as a group, did not discriminate its heartbeats.
Two things worth recording from that:
- d’ is the right score for the “manufacturable” worry, and does not resolve it. Unlike counting % error or a raw hit rate, d’ separates sensitivity from response bias, so the Nentjes result cannot be dismissed as offenders simply answering “simultaneous” more often. But a sample mean of exactly zero means the individual-difference signal is variation around chance, with roughly half the sample at negative d’ (systematic mis-timing, not weaker perception) — which is harder to read as “less of a real percept” than the counting literature’s low scores. The cardiodynamic and engagement confounds are untouched: stroke volume was not measured, and the offenders found the task “tedious.”
- A chance-level group can still carry external-criterion signal. The residual variance in d’ tracked psychopathy’s antisocial factor (Factor 2, r=−.29) even though the mean was at chance. Either the residual variance is a genuine (if small) perceptual signal the group mean hides, or it is being driven by something non-perceptual that also covaries with Factor 2 (the unexplained negative IQ→d’ effect in that paper, tedium/engagement, medication). The wiki cannot adjudicate this on one study, but it is the same “what is the score measuring in the people who cannot do the task?” question the validity debate asks of counting, arriving now on the discrimination side. See nentjes-2013-psychopathy-interoception.
The discrimination task, read across the lifespan — and a third confound it exposes
Khalsa, Rudrauf & Tranel (2009) run the discrimination family’s third variant on this page (a Brener–Liu–Ring simultaneity method: tones ~250–300 ms vs ~650–700 ms post R-wave, forced-choice, scored A’) across 59 adults aged 22–63, and find age accounts for 30% of the variance in accuracy — the steepest single demographic effect the wiki holds for any heartbeat task. Two things it contributes here:
- A tighter attention control than the counting literature runs. Khalsa et al. precede heartbeat detection with a pulse-detection control — the identical simultaneity judgment made against the participant’s own wrist pulse (an exteroceptive signal) — and exclude those who fail it. Pulse-detection accuracy does not predict heartbeat accuracy while age does, so the age decline is not just age-related decline in the attention/decision machinery a simultaneity task demands. Worth keeping as a template: an exteroceptive version of the same judgment is a cleaner attention control than a generic neuropsychological attention test.
- A cardiodynamic confound the age design makes vivid, and a mechanism datum that half-defuses it. Aging dampens cardiac rate and contractility, so an older heart is a quieter signal — and the cardiodynamic confound (below) then predicts Khalsa’s age decline with no loss of perceptual skill (stroke volume unmeasured, as ever). But Khalsa also supplies the literature’s sharpest evidence about what the felt heartbeat even is: cardiac transplant patients, before reinnervation, detect their heartbeats within the normal range (Barsky et al. 1998). So heartbeat “detection” does not primarily read cardiac afferent nerves — it reads something downstream (arterial pulsation at the skin; Pacinian fingertip sensitivity predicts detection, Knapp et al. 1997, and declines with age). This does not remove the loudness confound — a quieter pulse is a smaller skin signal too — but it does relocate it: the confound is about mechanical signal amplitude at whatever peripheral receptor does the work, not about cardiac innervation, and the aging effect is consistent with cutaneous receptor decline as much as with reduced cardiac output or cortical thinning. The task’s own author-lineage cannot say which, and neither can the wiki.
Role in this cluster
In Seth (2013) the task is the workhorse linking interoception to emotion, self, and body ownership:
- Interoceptive sensitivity measured this way correlates with right-AIC structure/function and with emotional symptomatology (Critchley et al. 2004).
- Lower scores predict greater rubber-hand-illusion susceptibility (Tsakiris et al. 2011) — central to the experience-of-body-ownership evidence.
- Modulates cardiac-timing effects on memory (Garfinkel et al. 2013).
Methodological caution
Seth’s 2013 usage of “interoceptive sensitivity” predates the now-standard Garfinkel et al. (2015) tripartite distinction (interoceptive accuracy = task performance; sensibility = self-reported/confidence; awareness = metacognitive correspondence). Heartbeat-counting in particular is critiqued for confounds with heart-rate beliefs. Interpret older “IS” claims accordingly.
The cardiodynamic confound
A distinct problem from the belief confound above, surfaced by Oldroyd et al. (2019) via Schandry et al. (1993) — the task’s own author, reporting the confound against his own instrument: stroke volume predicts heartbeat-detection performance — the more blood the heart pumps per beat, the better people estimate their own heart rate. The signal to be detected therefore varies in strength between people and within a person across states, and anything that raises sympathetic outflow (stress, hpa-axis activation, epinephrine) makes the task easier by making the heartbeat louder rather than the perceiver better. Chronically increased sympathetic outflow has even been proposed as a route to high interoceptive accuracy (Paulus & Stein 2010).
This cuts against reading heartbeat-detection differences as differences in perceptual skill, and it is not fixed by switching from counting to discrimination variants. It also complicates any group comparison where the groups plausibly differ in autonomic tone — which includes most clinical and stress-related comparisons the task is used for.
A near-miss test worth recording. Zoellner & Craske (1999) set out to test the arousal→accuracy link directly, inducing arousal with 600 mg of caffeine and finding accuracy flat as arousal rose — which reads at first like evidence against the confound. But their caffeine raised skin conductance, EMG, expired CO₂ and skin temperature while leaving heart rate unchanged (no trial effect for ECG). The cardiodynamic confound is about the cardiac signal’s amplitude; a manipulation that never moves the heart cannot test it, and their null holds only for non-cardiac arousal — which the confound never predicted a relationship with. If anything it points the confound’s way: the one channel that stayed flat (cardiac) is the one accuracy is meant to track, and accuracy stayed flat too. The manipulation that did move heart rate — Antony et al.’s exercise, above — is the one where accuracy followed. Zoellner & Craske ran the right experiment on the wrong variable.
What the score is for: a moderator, not a predictor
The usage this wiki should treat as the task’s most defensible one, from Dunn et al. (2010).
Across two studies (n = 58, n = 92), the Schandry score predicted nothing as a main effect — not felt arousal, not felt valence, not cardiac response to emotional images, not intuitive decision quality (r = .08, p = .46), not even how well the participant’s own body differentiated good options from bad (r = .07, p = .56). Everything it did, it did in a product term: it set how tightly bodily responses coupled to feeling and to choice.
That is a specific claim about what this task measures. It is not a measure of an ability that confers benefits; it is a measure of a channel’s gain. A high score means the body’s signal arrives loudly, and says nothing about whether the signal is worth hearing — the score does not even correlate with the signal’s usefulness. This is worth holding against the frontmatter strength that “individual differences predict AIC activation/morphometry and emotional symptoms”: those are correlations with the substrate and with symptoms, not evidence that a high score is good for anything.
And it complicates the cardiodynamic confound below rather than escaping it. If stroke volume drives both the score and the size of the cardiac signal available for coupling, then a “better perceiver couples more tightly” result is exactly what a pure signal-strength account predicts, with no perception involved. Dunn et al. report the interaction survives controls for Schandry confounds, but the list lives in an online supplement the wiki does not have. Recorded as unresolved.
One caution about generalizing the moderator reading. It is Dunn’s unselected design that yields “no main effect, only a moderation.” The counting task’s other decision-making application — Werner et al. (2009), good vs poor counters on the IGT — is extreme-groups and finds a straightforward main effect (good perceivers choose more advantageously, r=±.30), with the marker never measured so the moderation cannot even be tested. So “the score is a moderator, not a predictor” is a claim licensed by the design that can see moderations; on the selected-sample design the same score behaves like a predictor. Which is the artefact — the selected main effect (inflation) or the unselected null (dilution) — is the validity question restated, not settled.
But the null is contested, and the wiki should say so
The section above is written as though Dunn et al. settled what this task predicts. Pollatos, Kirsch & Schandry (2005) is the reason it did not, and this page carried the strong version for one ingest too long.
Pollatos et al. ran nearly the same experiment as Dunn et al.’s Study 1 — affective pictures for 6 s, SAM valence and arousal per picture, Schandry counting, healthy adults, n = 44 against 58 — and found the main effect Dunn reports as absent: good heartbeat perceivers rated affective pictures as more arousing, F(1, 39) = 5.90, P < .05, with r = 0.34 between the score and mean arousal. Their own introduction says this was the field’s consensus, citing Schandry (1981), Wiens et al. (2000), Critchley et al. (2004), Ferguson & Katkin (1996) and Montoya et al. (1993), against one contrary result (Blascovich et al. 1992).
Two things stop this from simply overturning the section above.
The samples differ in a way that predicts the gap. Pollatos et al. used extreme groups — 22 good perceivers selected from a screen of ~140, plus 22 age/sex-matched comparisons. That design has more power to detect a group difference than an unselected sample, and it inflates correlations computed across it. Their r = 0.34 and Dunn’s r = .08 are not two estimates of the same quantity, and the honest reading is that neither is a population value.
The two papers fit different models. Dunn’s claim is that accuracy is a gain term on the bodily signal, which requires measuring the bodily signal. Pollatos et al. recorded ECG and never analysed it against the pictures — so their design cannot evaluate Dunn’s model, and Dunn’s model arguably predicts their result: with no accuracy main effect but a real interaction, averaging over bodily responses that are not centred on zero leaves a marginal accuracy slope. (Dunn’s own marginal was null, which that account still owes an explanation for. Recorded on pollatos-2005-interoceptive-awareness-erp as this wiki’s reading, not either paper’s.)
A third explanation, added with the Van der Does et al. (2000) ingest, pointing opposite to the first. The wiki has been reading Pollatos’s extreme-groups design purely as a weakness — selection inflates correlations, so discount their r. If the task is valid only for the ~8–17% who are genuinely accurate, that same selection looks different: Pollatos screened ~140 people and took 22 good perceivers (≈16%), which is almost exactly the prevalence of genuine accuracy in the pooled data. Selecting the top sixth may be approximately selecting the subpopulation the task works on — purification, not just distortion. And Dunn’s unselected correlation then runs across a sample where roughly four in five participants contribute noise on the moderator, which attenuates toward zero and is the shape of r = .08.
Both readings can be true at once: selection can isolate a real subgroup and inflate the correlation within it. The wiki now records both, and notes the price the mixture reading charges Pollatos too — if accuracy is state-dependent and manufacturable by heart rate, “good perceiver” is not a stable category to select on, and selecting on one session to use as a trait moderator later is a problem neither study addresses. This is the wiki’s inference; Van der Does et al. predate Dunn by a decade and discuss neither study.
So what does the score predict? Unsettled. The defensible statement is narrower than either paper’s: across two studies using the same paradigm, the task’s relationship to felt arousal is design-dependent, and its relationship to felt valence is reliably nil (Pollatos: F(1, 39) = 0.14; Dunn: delta-R-squared = .00). The valence null is the part that replicates.
And note what both papers share, which is the deeper problem: neither can rule out cardiodynamics. A louder heart is easier to count, is a bigger signal to couple to, and plausibly comes with higher tonic arousal. That single account predicts Pollatos’s main effect, Dunn’s interaction, and Pollatos’s emotion-general P300 difference — with no perceiving anywhere in it. See below.
Its neural correlate is cardiac too (Haruki & Ogawa 2023)
One of the strongest arguments for taking this task seriously has been that its scores track right anterior insula activation — Critchley et al. (2004), Pollatos et al. (2007), Caseras et al. (2013) — placing it on the cortex Craig identifies as the substrate of interoceptive awareness. A behavioural score whose validity is otherwise contested gains a lot from converging with the anatomy.
Haruki & Ogawa (2023) narrow what that convergence licenses. When cardiac and gastric attention are compared directly in the same 31 people, the right dorsal anterior insula is the only region preferring the heart; and the accuracy-to-right-AIC relationship does not appear for awareness of breathing (Wang et al. 2019) or of skin conductance (Baltazar et al. 2021). Schulz’s (2016) supporting meta-analysis is, by its own title, of heart-focused interoception.
So the neural validation of this task is a cardiac validation, and it does not transfer to the construct the task’s scores are routinely used to stand for. This is the anatomical version of the point Banellis et al. made behaviourally and is-interoception-domain-general now hosts: a heartbeat score is a claim about a heart, at the behavioural level and at the cortical one.
It cuts one useful way for this page, though. The rival cardiodynamic reading above — that the score indexes how loud your heart is rather than how well you perceive it — gains nothing from this and arguably loses a little: a signal-amplitude account has no particular reason to predict that the anterior insula specifically prefers cardiac attention in a task where nothing is counted and no heartbeat is louder than another.
A third response format: motor tracking, and its built-in exteroceptive control
Counting (Schandry) and discrimination (Katkin/Brener) are the two branches above. García-Cordero et al. (2017) use a motor-tracking variant — tapping a key in time with each heartbeat — scored as an accuracy index against ECG. Two features earn it a note here. First, it ships with a matched exteroceptive control: the same tapping response made to an external audio heartbeat, which is the clean comparison this page keeps wishing the counting studies had (an exteroceptive version of the identical judgment, as Khalsa et al. built with pulse-detection). In that study, tracking an external heartbeat (M=0.73) beat tracking one’s own (basal M=0.47), the direction the loudness account predicts. Second, it makes interoceptive learning measurable within a session: one block of stethoscope feedback lifted own-heartbeat accuracy to M=0.58, and the gain showed up in HEP connectivity rather than HEP amplitude. The tradeoff is a motor component the counting task lacks — though the interoceptive-vs-exteroceptive contrast is between two tapping conditions, so it is controlled rather than confounding.
The one self-report it does correlate with (Murphy et al. 2020)
Nearly every section of this page records the score failing to relate to something. Here is the exception, and its shape is informative.
In two samples, counting accuracy correlated with the IAS — a trait questionnaire asking how accurately one perceives bodily signals — at r = .271 (n = 52) and r = .336 (n = 33), while correlating with the attention-flavoured BPQ at r ≤ .13 (ns) in the same participants and sessions, and with time-estimation performance, the ICQ and the TAS-20 not at all. Post-task confidence behaved identically. So the task is not simply orthogonal to self-report; it is orthogonal to self-reported attention and modestly related to self-reported accuracy.
That is a point for validity — a pure guessing-from-believed-heart-rate account has to explain the selectivity — and it is a small one: p ≈ .047 in both, ~7–11% shared variance, and the paper’s own control-variable analyses failed to replicate, so “survives controls for heart-rate beliefs” means the controls found nothing to remove. It also has a deflationary reading this page is well placed to supply: if a forceful heart both raises the count and teaches its owner that they can tell when it races, the cardiodynamic account predicts the correlation and its selectivity, with no perception in it. Stroke volume was not measured. See is-the-heartbeat-counting-task-valid.
Sibling measures
The task is the wiki’s only objective interoception measure, and its cost is narrowness: cardiac only, adult-friendly only, and silent on how a person relates to their body. The two instruments that enter the wiki with Oldroyd et al. trade objectivity for that missing coverage — the maia (self-reported sensibility across eight dimensions) and self-report-physiology-congruence (subjective/physiological correspondence, usable in children). The IAS and BPQ occupy the two self-report cells the section above separates. None substitutes for the others; see interoceptive-taxonomy.