Why This Matters Before You Trust a Pooled Number
When a meta-analysis combines results from several studies, the headline number — the “pooled effect” — can hide real disagreement between the studies feeding into it. That disagreement is called heterogeneity, and checking for it is one of the most important steps in deciding how much weight a meta-analysis result deserves. This guide explains what heterogeneity is, how researchers measure it, and what questions you can ask before treating a combined result as settled.
By ClinicalStudyConnect.com Research Desk | Updated September 2026
What “Heterogeneity” Actually Means
A meta-analysis is a statistical method for combining the numerical results of two or more separate studies into a single overall estimate. The idea is that pooling data improves precision and can settle disagreements between studies that appear to conflict.
But studies included in a meta-analysis are never identical. The Cochrane Handbook, the primary methods reference used by systematic reviewers worldwide, distinguishes between three overlapping types of variation:
- Clinical diversity — differences in who was studied, what exact intervention or dose was used, and which outcomes were measured.
- Methodological diversity — differences in study design, how outcomes were measured, and risk of bias between studies.
- Statistical heterogeneity — the observed intervention effects varying more than would be expected from chance alone, as a consequence of clinical diversity, methodological diversity, or both.
When researchers say a meta-analysis “shows heterogeneity,” they generally mean this last, statistical sense: the individual study results disagree with each other by more than random variation would predict.
The Evidence Ladder: How Heterogeneity Is Checked
Researchers do not simply eyeball whether studies agree. There is a standard sequence of checks, each adding more information than the last.
Step 1: Look at the Confidence Intervals
Each study in a meta-analysis reports an effect estimate along with a confidence interval — a range indicating the uncertainty around that estimate. If the confidence intervals from different studies barely overlap, that is a visual sign of possible heterogeneity.
Step 2: The Chi-Squared Test
A formal statistical test (often called the Chi2 or Cochran’s Q test) asks whether the differences between study results are compatible with chance alone. A low p-value suggests the differences are not simply random noise. This test has an important weakness: it has low statistical power when a meta-analysis includes few studies or small studies, which is common. That means a non-significant result does not prove there is no heterogeneity — it may simply mean there was not enough data to detect it.
Step 3: The I-Squared Statistic
Because the Chi-squared test only asks whether heterogeneity exists, researchers also use the I2 statistic to describe how much of the variability across studies is due to real differences rather than chance, expressed as a percentage. The Cochrane Handbook offers a rough, deliberately overlapping guide to interpretation:
- 0% to 40%: might not be important
- 30% to 60%: may represent moderate heterogeneity
- 50% to 90%: may represent substantial heterogeneity
- 75% to 100%: considerable heterogeneity
These ranges deliberately overlap because a given I2 value means different things depending on the size and direction of the effects involved and how many studies were pooled. A high I2 from only two or three studies is far less reliable than the same value from twenty studies.
Step 4: Prediction Intervals
A confidence interval around the pooled result tells you about the average effect. It does not tell you how much the true effect might vary from one study population to the next. A prediction interval addresses that by estimating the range within which the effect in a new, similar study would likely fall. Cochrane guidance recommends using prediction intervals only when a reasonable number of studies — generally five or more — are included, since they can look artificially wide or narrow with fewer.
What Researchers Do When Heterogeneity Shows Up
Finding heterogeneity does not automatically invalidate a meta-analysis, but it does change how the result should be handled. According to the Cochrane Handbook, the standard options include:
- Re-check the data. Apparent heterogeneity sometimes traces back to a data-entry or extraction error in one of the pooled studies.
- Decide not to pool at all. A systematic review does not have to include a meta-analysis. If results point in genuinely different directions, quoting a single average can be misleading.
- Investigate the cause. Subgroup analysis or meta-regression can explore whether a specific factor — dose, population, follow-up length — explains the disagreement, though this exploratory work should ideally be planned before looking at the results, not after.
- Use a random-effects model. A fixed-effect model assumes every study is estimating exactly the same true effect. A random-effects model instead assumes the true effect varies somewhat across studies and produces a wider, more conservative confidence interval when heterogeneity is present.
- Reconsider the effect measure. Sometimes apparent heterogeneity is an artifact of the statistic chosen (for example, using a raw mean difference when studies used different measurement scales) rather than a true disagreement.
- Examine outlier studies with caution. Removing a study because its result looks unusual is only justified when there is a clear, pre-identifiable reason for the difference — not simply because it changes the pooled result.
What’s Established vs. What Remains Uncertain
- Established: Some degree of heterogeneity is expected in nearly every meta-analysis, because no two studies use identical populations, doses, or methods.
- Established: A random-effects model and a fixed-effect model give identical results when there is no heterogeneity, and diverge as heterogeneity increases.
- Established: The choice between fixed-effect and random-effects models should never be based solely on the result of a statistical significance test for heterogeneity.
- Still debated: Statisticians continue to disagree over which method for estimating between-study variance performs best across different scenarios; no single approach is considered universally superior.
- Still limited: With very few included studies, both heterogeneity statistics (I2, Tau2) and the confidence intervals built around them are estimated poorly, and any interpretation should be treated cautiously regardless of which number comes out.
A Practical Checklist Before You Trust a Pooled Result
- Does the article report an I2 value or a heterogeneity test, and if so, what did it find?
- How many studies were pooled? Heterogeneity statistics from two or three studies are far less reliable than from a dozen or more.
- Did the researchers use a fixed-effect or random-effects model, and does that choice make sense given what was found?
- If heterogeneity was present, did the authors explore why — or did they report the pooled number without comment?
- Was a prediction interval reported alongside the confidence interval? If so, how wide is the range of plausible true effects?
- Did the review follow a structured reporting framework, such as the PRISMA guideline for systematic reviews, which promotes transparency about how studies were searched, selected, and combined?
Related Guides for Evaluating Pooled Research
For background on how meta-analyses fit alongside other research types, see our guide to Understanding Clinical Study Design Types. To learn how to weigh the overall strength of a body of evidence beyond a single statistic, see Evaluating Evidence Quality. For a closer look at how adverse-event data is handled across pooled studies, see Clinical Study Safety Data.
Medical Disclaimer
This article is for educational purposes only and does not constitute medical advice. It does not recommend, endorse, or advise for or against any product, treatment, or course of action. If you have questions about a specific health condition or treatment decision, consult a qualified healthcare provider. If you are experiencing a medical emergency, contact your local emergency services immediately.
Sources: Cochrane Handbook for Systematic Reviews of Interventions, Chapter 10: Analysing data and undertaking meta-analyses; PRISMA Statement (prisma-statement.org).