Skip to content
All guides

Is my I² too high?

Models & heterogeneity · Updated July 2026

Short answer

Probably not — and the question itself is slightly wrong. I² does not measure how much the true effects differ; it measures what proportion of the observed variation is due to real differences rather than chance. A set of very large trials can show a high I² over differences too small to matter, and a set of small trials can show a low I² while hiding differences that matter a great deal. Look at the between-study variance and the prediction interval instead.

What I² actually is

I² is a ratio: the share of the total variability in your effect estimates that is attributable to real between-study differences rather than to sampling error. It is a proportion, not a quantity. It has no units, and it is not on the scale of your outcome.

This has a consequence that catches experienced reviewers. Because I² is a share of the total variability, it rises as the within-study error falls. Bigger, more precise studies shrink the denominator. So the same real, trivial difference between true effects will produce a low I² in a set of small trials and a high I² in a set of large ones.

I² can approach 90% across a set of enormous trials whose true effects differ by an amount no clinician would act on. It can sit near 0% across a handful of tiny trials that are too imprecise to reveal the disagreement between them.

Why the thresholds mislead

The familiar bands — roughly 0–40% "might not be important", 30–60% "moderate", 50–90% "substantial", 75–100% "considerable" — are widely quoted as if they were decision rules. They deliberately overlap, and they are published as a rough guide with the explicit caveat that the importance of the observed value depends on the size and direction of the effects and on the strength of the evidence for heterogeneity.

Reading "I² = 62%, therefore substantial heterogeneity, therefore the review is unreliable" is a chain of three unearned inferences. Nothing about a percentage tells you whether the effects differ in a way that changes a decision.

What to look at instead

The between-study variance (τ²), and τ on the outcome scale
Unlike I², this is in the units of your effect measure. It answers "by how much do the true effects actually differ?" — which is the question you meant to ask.
The prediction interval
The range a new study might plausibly fall in. If the confidence interval sits entirely on one side of no-effect but the prediction interval crosses it, that is the single most useful thing you can tell a reader, and I² will not tell them.
The forest plot itself
Look at it. One outlying study driving the spread is a different situation from a smooth fan of disagreement, and both can produce the same I².
Pre-specified subgroups and meta-regression
If you have a hypothesis about why effects differ — dose, population, follow-up — test it. Explained heterogeneity is a finding. Unexplained heterogeneity is a limitation.

So what do I write in the paper?

Report I², because readers and reviewers expect it, but do not let it carry the argument. Report τ² and the prediction interval next to it, say whether the spread of effects is clinically important for your question, and say whether you could explain it. If heterogeneity is unexplained and material, that belongs in the certainty rating for the outcome — not in a sentence apologising for the percentage.

Where this answer stops

Everything here is about interpretation, not a rule you can automate. Whether the spread of effects matters is a clinical judgement about your question, and no statistic will make it for you.