The ROBINS-I tool assesses the risk of bias in non-randomised studies of interventions. Its defining idea is the target trial: the assessor imagines the hypothetical randomised trial the study is trying to emulate, then judges each domain by how far the real study falls short of that ideal, including the confounding that randomisation would have removed.
Why non-randomised studies need a different tool
When participants are not randomised, the groups being compared may differ in ways that affect the outcome before any intervention is given. That confounding is the central threat the RoB 2 tool never has to confront, because randomisation handles it by design. ROBINS-I exists precisely to make that threat assessable, which is why it sits beside RoB 2 in a complete risk of bias assessment whenever a review includes observational evidence.
The seven domains of ROBINS-I
Pre-intervention domains
Two domains are judged before the intervention begins. The first is bias due to confounding, the most important domain, which asks whether prognostic factors that predict the outcome also predicted who received the intervention. The second is bias in selection of participants into the study, where selection linked to both intervention and outcome can distort the result.
At-intervention domain
The third domain, bias in classification of interventions, asks whether intervention status was defined and recorded the same way for everyone, free of any knowledge of the outcome. Misclassification that depends on the outcome is the concern here.
Post-intervention domains
The remaining four domains cover deviations from intended interventions, missing data, measurement of the outcome, and selection of the reported result. The last three echo concerns you will recognise from RoB 2, but they are judged against the target trial rather than against an ideal randomised protocol alone.
Pre-specifying confounders before you score
ROBINS-I cannot be applied cold. Before assessing any study, the review team lists the important confounders and co-interventions for the question, ideally in the review protocol. Each study is then judged on whether it measured and adjusted for those named factors. Skipping this step turns the confounding domain into guesswork and is a common reason ROBINS-I assessments fall apart at peer review.
A concrete routine keeps this honest. Convene the clinical and methodological members of the team and agree, in writing, the handful of prognostic factors that plausibly predict both who received the intervention and the outcome, for example age, disease severity, and prior treatment for a study of a surgical technique. Add the co-interventions that could differ between groups and affect the outcome. A study that adjusted for all of the named confounders with a valid method can reach low or moderate risk on the confounding domain; one that adjusted for none, or used an unreliable method, cannot rise above serious however large its sample. Tying these named factors to the columns of the data extraction form means the information needed to score is captured as the study is read, not hunted for afterwards.
The seven-step procedure for one study
For each study and each outcome, two assessors should work independently through a fixed order before reconciling:
- State the target trial the study is emulating: the population, the intervention and comparator, and the outcome.
- Check the study against the pre-specified confounders and co-interventions agreed in the protocol.
- Answer the signalling questions for each of the seven domains, recording the supporting evidence.
- Reach a domain judgement of low, moderate, serious, or critical risk.
- Carry the worst domain forward, since the overall rating is governed by the most compromised domain, not an average.
- Decide whether a study at critical risk should be excluded from the synthesis or reported separately.
- Compare the two independent assessments and resolve disagreements by discussion or a third assessor.
From judgement to certainty
ROBINS-I rates each domain on a four-level scale of low, moderate, serious, or critical risk, the last reserved for studies too compromised to include in a synthesis. Because non-randomised evidence starts lower in the GRADE certainty framework, a clean ROBINS-I result is what lets a review argue that observational studies still deserve a credible certainty of evidence rating rather than being dismissed for their design alone. The judgements are reported alongside the extracted outcome data so readers can see why each study was trusted or set aside, and they feed directly into the wider appraisal record for the review.
How ROBINS-I compares with simpler tools
Teams sometimes default to a lighter instrument such as the star-based Newcastle-Ottawa Scale for cohort and case-control evidence, and for a descriptive review that can be defensible. ROBINS-I earns its extra weight when the review makes a causal claim from observational data, because only the target-trial logic forces an explicit account of confounding and a comparison against a defined ideal. Whichever tool you choose should be named before the results are seen, exactly as you would fix the randomised-trial tool for the trials in the same review, so the appraisal cannot be tuned to the conclusion. If applying ROBINS-I across a mixed body of evidence is more than your team can carry, our risk of bias appraisal service runs it with assessors trained on the target-trial approach.
Common mistakes when applying ROBINS-I
The recurring failures are concentrated in the confounding domain. Assessors score without a confounder list, so the most important domain becomes an opinion; they credit statistical adjustment uncritically, when an adjustment for the wrong variables or with an invalid model does not actually control bias; and they average across domains, letting four tidy domains outvote a serious confounding problem when the overall rating should follow the worst domain. A final, subtler error is forgetting to define the target trial at all, which leaves the whole assessment without the reference point it depends on.