Two different things get called “cutoffs” for this instrument, and conflating them is the source of most of the confusion. One kind is supported by published validation work. The other is not.
The instrument
The Zarit Burden Interview began with Zarit, Reever and Bach-Peterson in The Gerontologist in 1980. The revised version in general use has 22 items, each scored 0 to 4, giving a total range of 0 to 88, with higher scores indicating greater perceived burden. Short forms exist with 13, 12, 7, 6, and 4 items.
A licensing note before anything else. The instrument is copyrighted and distributed under license through the Mapi Research Trust (ePROVIDE). Permission is required, including for research use. This page discusses how scores are interpreted; it does not reproduce the items.
Screening cutoffs: these are real
Individual validation studies have derived thresholds for identifying caregivers who warrant further assessment. On the 22-item version, those cluster in the mid-twenties: a curated rehabilitation measures database reports cutoffs ranging from 24 to 26, and notes that a cutoff of 25 correctly identified 77 percent of high-burden stroke caregivers.
For short forms, a validation study in 394 dementia caregivers found the 6-, 7-, and 12-item versions performed comparably to the full instrument, with sensitivity around 77 to 85 percent and specificity around 60 to 80 percent. It recommended the 6-item version as the practical choice, with a cutoff of 9 or above on its 0 to 24 range.
Two qualifications belong with those numbers. They were validated in particular populations — stroke caregivers, dementia caregivers — and do not automatically transfer. And several were validated against caregiver depression as the reference standard rather than against an independent measure of burden, which makes them closer to depression-screening thresholds than to burden grades.
Severity bands: these cannot be traced
A four-level scheme circulates widely: 0–21 little or no burden, 21–40 mild to moderate, 41–60 moderate to severe, 61–88 severe. It appears on clinical websites, in review articles, and in the discussion sections of trials.
We could not trace it to any primary validation study. The citation trail leads to secondary sources citing other secondary sources. Note also that the boundaries overlap — a score of 21 falls in two bands at once, and a scheme derived from data would not do that. Overlapping boundaries are the fingerprint of a table that has been copied rather than calculated.
What curated instrument databases do is instructive. One lists the ZBI with no interpretation ranges at all, describing only the structure and pointing readers to the licensor. Another reports the study-specific screening cutoffs above — and no severity bands. The people who catalogue these instruments for a living are not publishing the four-band scheme.
Population means, which are more useful than bands
Reported mean scores differ substantially by the condition being cared for. Comparing a score against the wrong reference population will mislead in a way no severity label captures.
| Caregiver population | Reported mean ZBI-22 |
|---|---|
| Stroke | 28.3 |
| Dementia | 24.4 |
| COPD | 20.4 |
Means as reported in a curated rehabilitation measures database. They describe samples, not thresholds.
What the systematic reviews say
A 2024 systematic review conducted under COSMIN guidelines examined short forms of the instrument across 66 articles. It found that most versions showed satisfactory content validity and high internal consistency — and also that measurement invariance, criterion validity, and test-retest reliability were not established for all the versions analyzed, and that structural validity was not satisfactory for all of them.
The dimensionality question remains genuinely open. Published factor analyses support one-, two-, three-, four-, and five-factor solutions. In a study of 194 dementia caregivers the 12-item version resolved into “personal strain” and “role strain,” and only personal strain predicted caregiver psychological distress. Separately, an item response theory analysis argued the 12-item version is unidimensional. Both cannot be right, and the practical consequence is the same either way: a total score can move on items that do not track the outcome anyone cares about.
What this means if you are reporting a change
There is no well-established minimal important change value for the Zarit Burden Interview. That is not a gap anyone hides; it simply has not been established. The consequence is direct: a trial reporting that an intervention produced a “clinically meaningful” improvement in burden owes the reader a threshold and a citation for where that threshold came from. Absent both, the claim cannot be evaluated.
This matters beyond methodology. The Cochrane review of remote and web-based interventions for dementia caregivers found little or no effect on burden — a standardized mean difference of −0.06, with a confidence interval from −0.35 to 0.23, across 26 trials and 2,367 participants. Anyone building in this area is measuring against a scale whose interpretive thresholds are weaker than the scale's ubiquity suggests. That is worth stating plainly rather than working around.
Common questions
What is the score range of the Zarit Burden Interview?
The revised 22-item Zarit Burden Interview scores each item 0 to 4, giving a total range of 0 to 88. Higher scores indicate greater perceived burden. Short forms with 13, 12, 7, 6, and 4 items also exist, each with its own range.
What is a high score on the Zarit Burden Interview?
There is no single answer, and that is the honest one. Validated screening cutoffs from individual studies cluster around 24 to 26 on the 22-item version; one study found a cutoff of 25 correctly identified 77 percent of high-burden stroke caregivers. These are screening thresholds validated in specific populations, not universal severity grades.
Are the 0-21, 21-40, 41-60, 61-88 severity bands valid?
They cannot be traced to a primary validation study. They appear on many clinical websites and in many papers, but the citation trail does not lead to an original source. The boundaries also overlap at 21, which is a sign of a scheme propagated by copying rather than derived from data. Curated instrument databases do not list them.
What are typical Zarit scores by caregiver population?
Reported means differ by the condition being cared for: approximately 28.3 for stroke caregivers, 24.4 for dementia caregivers, and 20.4 for COPD caregivers. Comparing a score against the wrong reference population will mislead.
What counts as a meaningful change in a Zarit score?
No well-established minimal important change value exists. A 2024 systematic review following COSMIN guidelines, covering 66 articles on short forms of the instrument, found that measurement invariance, criterion validity, and test-retest reliability were not established for all versions, and that structural validity was not satisfactory for all versions. Any claim that an intervention produced a clinically meaningful improvement on this scale should say what threshold it is using and where that threshold came from.
Can I use the Zarit Burden Interview for free?
No. The instrument is copyrighted and distributed under license through the Mapi Research Trust (ePROVIDE). Permission is required, including for research use. This page describes score interpretation and does not reproduce the instrument.
Sources
- Zarit S.H., Reever K.E. & Bach-Peterson J., “Relatives of the impaired elderly: correlates of feelings of burden,” The Gerontologist 20(6):649–655, 1980.
- Bédard M. et al., the Zarit Burden Interview short form, The Gerontologist 41(5):652, 2001. Oxford Academic.
- Yu J., Yap P. & Liew T.M., comparison of Zarit Burden Interview short versions in 394 dementia caregivers, Aging & Mental Health 23(6):706, 2019. Taylor & Francis.
- Cejalvo E., Gisbert-Pérez J., Martí-Vilar M. & Badenes-Ribera L., “Systematic review following COSMIN guidelines: short forms of Zarit Burden Interview,” Geriatric Nursing 59:278–295, September–October 2024. PubMed 39094351.
- Internal consistency, sensitivity, specificity and predictive value of the ZBI-7: systematic review and generalization meta-analysis, Journal of Affective Disorders, 2025. PubMed 40555351.
- Shirley Ryan AbilityLab, Rehabilitation Measures Database, Zarit Burden Interview entry, for screening cutoffs and population means. sralab.org.
- EdInstruments, Zarit Burden Inventory entry, which lists no interpretation ranges. edinstruments.org.
- Mapi Research Trust / ePROVIDE, Zarit Burden Interview licensing record. ePROVIDE.
- Review of the psychometric properties of the Zarit Burden Interview, Frontiers in Psychology, 2022. Frontiers.
- Factor structure of the ZBI-12 and prediction of caregiver distress, Journal of Applied Gerontology, 2014. SAGE.
- Unidimensional 12-item Zarit Caregiver Burden Interview obtained by item response theory, Value in Health. Value in Health.
- Cochrane Database of Systematic Reviews, remote and web-based interventions for dementia caregivers, CD006440.pub3.
Related
What the evidence shows about technology for dementia caregiversThe Cochrane review found little or no effect on burden. This is the scale it was measured on. Documentation in home-based care: what is measured, and what isn’tThe same problem on the clinical side: one federal number, and very little else measured.Published July 31, 2026. Corrections to rupakshah@pyramid.care.