What does GRADE certainty actually mean?
GRADE certainty is your confidence that the estimated effect is close enough to the truth to support a decision. It is rated per outcome, not per study and not for the review as a whole. "Low certainty" does not mean the effect is small or that the studies were bad — it means that further research could well change the estimate.
The four levels, in plain terms
- High
- We are very confident the true effect lies close to the estimate. Further research is very unlikely to change it.
- Moderate
- We are moderately confident. The true effect is probably close, but there is a possibility it is substantially different.
- Low
- Our confidence is limited. The true effect may be substantially different from the estimate.
- Very low
- We have very little confidence. The true effect is likely to be substantially different.
Note what every level is about: the estimate, and how much it might move. None of them is a statement about whether the intervention works, or about how good the trials were in the abstract.
What moves the rating
Randomised trials start at high certainty and observational evidence starts at low. From there, five considerations can lower the rating: risk of bias in the contributing studies, inconsistency between them, indirectness of the population or outcome relative to your question, imprecision in the pooled estimate, and suspicion of publication bias.
Three considerations can raise it, and they apply in practice only to observational evidence: a large magnitude of effect, a dose-response gradient, and the situation where plausible confounding would have worked against the effect that was nonetheless observed.
The three misreadings to avoid
- "Low certainty means it does not work"
- It means you do not know. Low certainty is compatible with a large true benefit and with no benefit at all — that is precisely the problem it is flagging.
- "The review is moderate certainty"
- A review does not have a certainty rating. Each outcome does, and they routinely differ: high certainty for a common outcome with many events, very low for a rare harm nobody powered for.
- "The trials were well conducted, so certainty is high"
- Risk of bias is one of five reasons to downgrade. Flawless trials that are too small, or that measured a surrogate rather than the thing you care about, still yield low certainty.
The rating is only as useful as the reasons attached to it. A summary of findings table that reports "Low" without saying which domains were downgraded and why is not GRADE — it is a label.
Where this answer stops
GRADE is a structured judgement, not a calculation. Two competent teams can rate the same evidence differently and both be defensible. Its value is that the reasoning is explicit and contestable, not that it is reproducible to the level of a statistic.