The RoB 2 tool is the Cochrane instrument for judging the risk of bias in a randomised controlled trial. It works through five fixed domains using a set of signalling questions, and translates the answers into a domain judgement of low risk, some concerns, or high risk, before rolling those up into an overall judgement for each outcome the trial reports.

What changed from the original tool

RoB 2 replaced the older domain-based approach with a far more prescriptive process. Instead of asking an assessor to form a holistic view, it routes them through structured signalling questions whose answers feed an algorithm that proposes the judgement. This reduces the room for an assessor’s prior beliefs to drive the result and makes two independent assessors far more likely to agree, which matters when you are scoring this in duplicate as part of a wider risk of bias assessment.

The five domains of RoB 2

1. Bias arising from the randomisation process

This domain asks whether the allocation sequence was truly random and adequately concealed, and whether baseline differences hint at a broken randomisation. A failure here is a form of selection bias that can bias the result from the very start.

2. Bias due to deviations from intended interventions

Here the assessor considers whether participants and carers knew their assigned group, and whether the analysis followed an intention-to-treat principle. Deviations that arise because people knew their assignment are the classic performance bias concern.

3. Bias due to missing outcome data

Loss to follow-up matters when it is related to the outcome. The signalling questions probe how much data were missing and whether that loss could plausibly depend on the true outcome value.

4. Bias in measurement of the outcome

This domain turns on whether outcome assessors were blinded and whether the method of measurement could differ between groups. Subjective outcomes rated by unblinded assessors are the most exposed, which is why the same trial can land at different risk levels for an objective lab value and a self-reported symptom score.

5. Bias in selection of the reported result

The last domain checks the result against a pre-specified analysis plan or registered protocol to catch selective reporting, where the most favourable of many analyses is the one presented. A protocol on PROSPERO or a trial registry is what makes this domain answerable.

The two analysis effects you must choose between

Domain 2 hides a decision that catches many assessors out: RoB 2 asks whether you are estimating the effect of assignment to the interventions (the intention-to-treat effect, what would happen if everyone were offered the treatment) or the effect of adhering to them (the per-protocol effect, what happens if everyone actually took it). The signalling questions, and therefore the rating, differ depending on which you choose. Most systematic reviews of effectiveness want the effect of assignment, so they judge whether the analysis respected intention-to-treat and whether deviations went beyond what would happen in routine care. Picking the effect of interest before you score, and stating it in the review protocol, is what stops the domain becoming a moving target.

From domain judgements to the overall rating

The overall judgement for an outcome is not an average. A trial is at high overall risk of bias if any single domain is high risk, and has some concerns if several domains raise concerns even when none is outright high. The algorithm proposes this verdict mechanically from the domain ratings, but it can be overridden with a recorded justification, for instance when a domain is technically “some concerns” yet the assessor judges the likely impact on this particular outcome to be negligible. Every override should be documented so the judgement stays auditable. Those judgements then carry into the GRADE certainty rating, where they inform whether the pooled forest plot estimate is downgraded for study limitations. For non-randomised designs you reach instead for the ROBINS-I target-trial tool, which is built around a different logic.

Scoring RoB 2 step by step

In practice, two assessors apply the tool independently to each outcome before reconciling. A workable sequence is:

  1. Fix the effect of interest (assignment or adherence) and the specific outcome being judged, since the same trial is rated separately for each outcome it reports.
  2. Answer every signalling question in a domain as “yes”, “probably yes”, “probably no”, “no”, or “no information”, citing the source text.
  3. Read off the domain judgement the algorithm proposes, and record a reasoned override only where it is warranted.
  4. Derive the overall judgement from the worst domain, then have the second assessor compare and resolve any disagreement by discussion or a third reviewer.

This is the same duplicate-and-reconcile discipline used in full-text screening, applied to appraisal, and it sits at the centre of any credible study-appraisal workflow. If running it across a whole review is more than your team can carry, our appraisal service applies it in duplicate with calibrated assessors.

Reporting RoB 2 in your review

Reviews present RoB 2 as a traffic-light table, one row per study and one column per domain, plus a weighted summary bar chart that shows the proportion of studies at each risk level. Recording the answer to each signalling question, not just the final colour, is what lets a reader audit the judgement and keeps you aligned with the PRISMA 2020 checklist.

Common mistakes when applying RoB 2

The errors that recur are predictable. Assessors rate the study once rather than once per outcome, erasing the fact that an objective laboratory value and a subjective symptom score from the same trial often deserve different verdicts. They skip the effect-of-interest choice and so answer Domain 2 inconsistently across studies. They treat “no information” as if it were “low risk”, when an unreported method is a reason for concern, not reassurance. And they override the algorithm silently, which breaks the reproducibility the tool was designed to deliver. Pre-specifying the approach and documenting every judgement removes all four.