Critical appraisal tools are structured instruments, checklists, scales, and domain-based frameworks, used in a systematic review to judge how trustworthy each included study is. The right tool is dictated by study design: Cochrane RoB 2 for randomised trials, ROBINS-I for non-randomised intervention studies, the Newcastle-Ottawa Scale for cohort and case-control studies, QUADAS-2 for diagnostic accuracy, AMSTAR 2 and ROBIS for systematic reviews themselves, CASP and JBI checklists across general and qualitative designs, and the Mixed Methods Appraisal Tool for mixed studies reviews, with GRADE sitting above them all to rate the certainty of a body of evidence.
How to choose a tool: the decision logic
Tool choice is not taste, and it is not a search for the “best” instrument in the abstract, because no such instrument exists. It is a two-question decision made at protocol stage. First, what is being appraised: a primary study, or an existing review? Second, what design is it: randomised, non-randomised interventional, observational aetiological, diagnostic, qualitative, or mixed? Answer those and the field narrows to one or two candidates per design, which is why every step-by-step account of the stages of a systematic review places tool selection before searching even begins. Two further constraints matter. Journals and Cochrane-style reviews increasingly mandate domain-based bias tools for intervention evidence, so check the venue’s expectations early. And every tool you name in the protocol must be applied exactly as named; swapping instruments after seeing the studies is a bias of its own.
One more distinction saves confusion later: appraisal tools are not reporting guidelines. Statements such as PRISMA 2020 for reporting a review or STROBE for reporting observational studies tell authors what to write down; appraisal tools judge what was actually done. A study can comply with every reporting item and still be at high risk of bias, so citing a reporting checklist as your appraisal instrument, a mistake that still appears in submitted manuscripts, appraises nothing at all.
Randomised trials: Cochrane RoB 2
For randomised controlled trials the default is the Cochrane RoB 2 instrument, which judges five bias domains per outcome through signalling questions and an algorithm. Its logic, domains, and common scoring errors are covered in our full guide to RoB 2; the point for tool selection is that RoB 2 judgements are the input GRADE expects for trial evidence, so choosing anything weaker for a trial-based review needs justification. Note that RoB 2 is applied per outcome, not per trial: the same study can be low risk for mortality and high risk for a self-reported symptom score, and reviews that rate each trial once are already deviating from the method. Variants exist for cluster-randomised and crossover trials, so check the design before assuming the standard form fits.
Non-randomised interventions: ROBINS-I
Where interventions were not randomised, cohort-style drug studies, policy evaluations, surgical series with comparators, the Cochrane family answer is ROBINS-I, built on target-trial logic: each study is judged against the hypothetical randomised trial it is trying to emulate, with confounding and selection into the study as the headline domains. It is the most demanding instrument in common use, and that is the price of taking non-randomised effect estimates seriously. Expect it to take substantially longer per study than any checklist, and expect many observational studies to land at serious or critical risk once confounding is judged honestly. Teams that find themselves reaching for a lighter tool to avoid that verdict should recognise the temptation for what it is.
Cohort and case-control studies: the Newcastle-Ottawa Scale
For observational aetiology and prognosis questions, many reviews still use the star-based Newcastle-Ottawa Scale, which awards up to nine stars across selection, comparability, and outcome or exposure. It is fast and familiar to journals, but its summed stars compress very different flaws into one number, so pair it with per-item reporting or consider a design-matched checklist instead. The comparability domain is where the scale earns or loses its keep: it is the only place confounding is examined, and reviews should state in the protocol which specific confounders a study must control to earn its stars, otherwise two appraisers will award them on different grounds.
Diagnostic accuracy: QUADAS-2
Diagnostic test accuracy studies have failure modes all their own, spectrum bias, unblinded index tests, partial verification, and no intervention-focused instrument asks about any of them, which is why diagnostic reviews get QUADAS-2, a tool built precisely for them: four domains judged for risk of bias, three of them also for applicability to the review question, with extensions (QUADAS-C, QUADAS-AI) for comparative and artificial-intelligence questions. Two features distinguish it from the intervention tools: the signalling questions are meant to be tailored to each review before use, and applicability is scored as its own axis rather than folded into bias, so a sound study of the wrong population is flagged as exactly that.
Systematic reviews themselves: AMSTAR 2 and ROBIS
When the unit being appraised is a review rather than a study, two instruments compete. AMSTAR 2 and its sixteen items rate confidence in a review’s conduct, with seven critical domains driving the verdict; ROBIS instead assesses risk of bias in the review process across three phases and four domains covering eligibility criteria, study identification and selection, data collection and appraisal, and synthesis. AMSTAR 2 is the more widely used and the easier to apply; ROBIS is the stricter bias instrument, and it is the one guideline programmes tend to require when a review will directly underpin recommendations. Both are the working tools of umbrella reviews of existing reviews, which stand or fall on how well they appraise what they include, and of the umbrella review service we run for teams synthesising review-level evidence.
General and qualitative appraisal: CASP
The Critical Appraisal Skills Programme publishes plain-language checklists across the common designs, each structured around validity, results, and applicability, as unpacked in our design-by-design CASP guide. CASP is the teaching-friendly entry point to appraisal, and its qualitative checklist is the dominant instrument in qualitative evidence synthesis, where formal bias tools do not reach.
A checklist per design: the JBI suite
The Joanna Briggs Institute maintains the broadest family of all, with design-matched checklists for randomised trials, quasi-experimental, cohort, case-control, cross-sectional, prevalence, case report, case series, qualitative, and text and opinion evidence, surveyed in our JBI appraisal guide. Its prevalence checklist in particular has no real competitor, which makes JBI the default for proportion meta-analyses and for reviews admitting designs the Cochrane tools ignore. All the checklists share one response format, yes, no, unclear, or not applicable, and end in an explicit decision to include, exclude, or seek further information, which keeps a multi-design appraisal legible in a single table.
Mixed methods evidence: the MMAT
Reviews combining qualitative and quantitative studies can appraise everything with one instrument, the Mixed Methods Appraisal Tool in its 2018 version: two screening questions, then five criteria matched to each of five design categories. It trades depth for coherence, one quality vocabulary across the whole review, and pairs well with a deeper design-specific tool for the component carrying the most inferential weight.
Above the studies: GRADE
GRADE is routinely listed alongside these tools and routinely misunderstood. It is not a study-level appraisal instrument at all: GRADE rates the certainty of evidence for each outcome across the whole body of studies, taking the study-level judgements from RoB 2, ROBINS-I, or QUADAS-2 as one input among five, alongside inconsistency, indirectness, imprecision, and suspected publication bias. Appraisal tools feed GRADE; they do not replace it, and a review needs both layers. The implication runs in both directions: an appraisal that produced only summary scores gives GRADE nothing usable, because the certainty framework needs to know which studies carry which limitations for which outcomes, while a GRADE exercise performed without study-level appraisal is guesswork wearing a framework. Plan the two together at protocol stage so the appraisal output arrives in the shape the certainty assessment consumes.
Risk of bias versus methodological quality
The families above split along a conceptual line worth respecting. Risk of bias tools (RoB 2, ROBINS-I, QUADAS-2, ROBIS) ask one question: could the study’s methods have systematically distorted its results? Methodological quality instruments (CASP, JBI, the Newcastle-Ottawa Scale, the Mixed Methods Appraisal Tool) ask more broadly whether the study was well conducted and well reported, which includes bias but is not limited to it. A study can be immaculately reported yet biased, and biased studies can be beautifully written. The distinction, and why modern tools abandoned numeric scores in favour of judgements, is treated at length in risk of bias versus quality assessment.
Common tool-selection mistakes
The recurring failures are worth listing because journals now reject for them. Using one generic checklist across every design in a multi-design review, which produces answers to the wrong questions for most of the studies. Applying a trial tool to observational evidence, or a quality checklist where the venue expects a formal bias instrument. Inventing a numeric threshold on a tool whose authors refused to publish one. Citing a reporting guideline as the appraisal method. Appraising with an outdated version of a revised instrument without saying so. And the quiet one: choosing the tool after reading the studies, when the shape of the evidence has already suggested which instrument will flatter it. Every one of these is prevented by the same habit, fixing the tool, version, tailoring, and decision rules in the protocol before the first study is appraised. Where a review includes designs from several families, name every instrument and the classification rule that assigns studies to them, because the assignment itself is a judgement two reviewers can disagree about, and an unstated rule is unauditable.
The appraisal workflow: always in duplicate
Whichever instrument you land on, the workflow is constant. Name the tool, version, and any tailoring in the protocol. Pilot it on three to five studies so both appraisers calibrate. Then appraise every study in duplicate, independently, citing the source text behind each judgement, and reconcile disagreements by discussion or a third reviewer, tracking agreement just as you would inter-rater reliability at screening. Finally, make the results consequential: feed them into GRADE, sensitivity analyses, and the discussion, with per-item tables in the supplement, the pattern described in our overview of risk of bias assessment. Teams short of a calibrated second appraiser can bring in our independent second reviewer support for exactly this stage. Choose by design, apply in duplicate, report per item, and let the appraisal change the conclusions: that is the whole discipline, whatever the tool, and it is the one part of a review no instrument can automate.