GRADE certainty of evidence is a structured way to rate how much confidence you can place in the pooled estimate for each outcome in a systematic review. It assigns one of four levels, high, moderate, low, or very low, by starting from the study design and then moving the rating up or down for specific, named reasons rather than an overall impression.

Why a separate certainty rating exists

A tight confidence interval on a forest plot can look convincing while resting on flawed studies. GRADE separates the precision of an estimate from the trust you should place in it. It answers a different question from the per-study risk of bias assessment: not “is this one study sound?” but “across all the evidence for this outcome, how sure can a reader be?”

Starting point and the five reasons to downgrade

Randomised trials start at high certainty and observational studies start at low. From there, GRADE can downgrade an outcome for five reasons.

1. Risk of bias

If the studies contributing to an outcome carry serious risk of bias, judged with RoB 2 or ROBINS-I, the certainty drops. This is the most direct link from study-level appraisal to the overall rating.

2. Inconsistency

Wide, unexplained heterogeneity between studies, with results pointing in different directions, lowers confidence that any single pooled number captures the truth.

3. Indirectness

When the included populations, interventions, or outcomes differ from the question actually being asked, the evidence is indirect and the rating falls.

4. Imprecision

If the confidence interval is so wide that it spans both meaningful benefit and meaningful harm, or rests on very few events, the outcome is downgraded for imprecision.

5. Publication bias

When small studies appear to be missing, the body of evidence may overstate the effect, so certainty is lowered. Reviewers probe this with a funnel plot and formal tests, as covered in our piece on publication bias.

When certainty can be upgraded

Observational evidence can move up for a large effect, a clear dose-response gradient, or when plausible confounding would only have reduced the observed effect. These upgrades are uncommon and require an explicit justification. A large effect upgrade conventionally applies when the relative effect measure is at least twofold (or halved) with no plausible confounder that could explain it, and a very large effect of around fivefold can justify two levels. The logic is that a residual confounder strong enough to manufacture such an effect would usually have been noticed. Upgrading is only ever available to evidence not already downgraded for the criteria above, so a body of observational studies at serious risk of bias cannot be rescued by a large-effect upgrade.

A worked rating: from design to final level

Take an outcome supported by four randomised trials. The body of evidence starts at high. Suppose two of the four trials had unblinded outcome assessors and lost a fifth of participants, judged with the RoB 2 signalling questions, so risk of bias is serious and you drop one level to moderate. The trials’ results then scatter widely with an I-squared above seventy-five per cent and no subgroup explaining it, which is unexplained inconsistency, dropping a second level to low. If the pooled confidence interval around the estimate also spanned both appreciable benefit and appreciable harm, that imprecision would drop a third level to very low. Each step is named and justified, so a reader can see precisely why a set of randomised trials ended up rated very low rather than being told to take the conclusion on trust.

Rating per outcome, not per review

GRADE is applied one outcome at a time, never as a single verdict on the whole review. The same set of studies can support a primary efficacy outcome at moderate certainty and a rare adverse event at very low, because the event count, the directness, and the contributing studies differ. Decide which outcomes matter before you rate, ideally the ones named in the review’s PICO question, and rate each on its own evidence. Collapsing several outcomes into one certainty level is a common error that hides exactly the distinctions a guideline panel needs.

Reporting GRADE in a Summary of Findings table

The output is a Summary of Findings table: one row per outcome, the pooled effect, the certainty level, and the reasons for any up- or downgrading. The standard columns are the absolute and relative effect, the number of participants and studies, the certainty symbol, and a footnoted reason for every move away from the starting level. Producing it well sits alongside the wider PRISMA 2020 reporting checklist and the extracted outcome data that the ratings draw on, so a reader can trace every judgement back to its source. If you would like the table built to a guideline-ready standard on top of a defensible appraisal, that is part of our certainty and appraisal service.

Who rates, and at what point in the review

GRADE is applied near the end of the review, once the studies are pooled and appraised, but the groundwork is laid much earlier. The outcomes to be rated should be chosen and ranked by importance at the protocol stage, before any results are seen, so the team cannot quietly promote whichever outcome happened to reach significance. The rating itself is best done by two reviewers working from the same evidence and reconciling, the same duplicate discipline used elsewhere in the review, because the downgrade decisions involve judgement and two people surface the borderline calls for discussion. A clinical or content expert is valuable for the indirectness and imprecision judgements, since deciding whether a confidence interval crosses a clinically important threshold needs to know what that threshold is for the question at hand. Recording who made each call, and on what basis, keeps the Summary of Findings table defensible when a guideline panel or an editor interrogates it line by line.

Common mistakes when applying GRADE

The most frequent error is double-counting: downgrading for risk of bias and again for imprecision when the same small, high-risk study is the real cause of both, which can push an outcome two levels for what is essentially one problem. A second is downgrading on impression rather than on a stated rule, for instance lowering for inconsistency without quantifying the heterogeneity behind it. A third is forgetting that certainty and precision are different things; a narrow interval on a pooled forest plot built from biased trials is precise but not certain. Stating the reason and the threshold for every step, in a footnote, is what keeps the rating auditable and stops it drifting to match a preferred conclusion.