AMSTAR 2 critical appraisal is a tool for judging the methodological soundness of a systematic review itself, not of the primary studies inside it. It works through sixteen items, several of which are flagged as critical domains, and produces an overall confidence rating of high, moderate, low, or critically low based on which weaknesses are present.

Appraising the review, not the studies

Most appraisal tools point downward at the included trials. AMSTAR 2 points at the review. It is what an umbrella review or a guideline team uses to decide whether an existing systematic review was conducted well enough to rely on. That makes it the natural companion to the risk of bias assessment a review runs on its own primary studies: one looks inward at the evidence, the other looks at the method that assembled it.

The critical domains that drive the rating

Not all sixteen items count equally. AMSTAR 2 designates several as critical, because a flaw in any one of them can undermine the whole review. These typically include whether the review used a pre-registered protocol, whether the literature search was adequate, whether excluded studies were justified, whether risk of bias was assessed in the primary studies, whether appropriate methods were used in any meta-analysis, and whether the impact of bias on the synthesis was considered.

How weaknesses map to confidence

A single critical flaw drops a review to low confidence, and more than one drops it to critically low, regardless of how well the non-critical items score. Weaknesses only in non-critical items leave a review at moderate confidence. This structure deliberately stops a review with a fatal search or analysis flaw from being rescued by tidy reporting elsewhere.

The seven critical items in plain terms

It helps to know which items carry the heaviest weight before you appraise anything. The seven domains AMSTAR 2 treats as critical are:

  1. A pre-registered protocol that names the research question and methods before the review began (item 2).
  2. An adequate, comprehensive literature search across multiple sources, not a single database (item 4).
  3. A justified list of excluded studies at the full-text stage, with reasons (item 7).
  4. A satisfactory risk of bias assessment of the included studies (item 9).
  5. Appropriate statistical methods for any pooling, such as a sensible model choice and weighting (item 11).
  6. Accounting for risk of bias when interpreting the results of the synthesis (item 13).
  7. An assessment of publication bias and a discussion of its likely impact (item 15).

The remaining nine items, covering things like the use of a question framework such as PICO or its variants, duplicate study selection, funding disclosure, and reporting of study characteristics, are non-critical. They still matter, but a weakness in them only nudges a review down to moderate, never to low or critically low on their own.

The items that catch most reviews out

Two non-obvious requirements trip up many authors. The first is justifying the list of excluded studies: reviews must present the studies that reached full text but were rejected, with reasons, which ties straight back to a clean PRISMA flow diagram. The second is that the review must investigate publication bias and discuss its likely effect, not merely mention it. A review that pools results but never probes for small-study effects will lose a critical item.

Two more catch out experienced teams. Item 9 asks for a satisfactory risk of bias appraisal, and reviewers routinely lose it by using a tool that does not match the designs, for instance a generic checklist where RoB 2 for randomised trials or ROBINS-I for non-randomised studies was needed. Item 11 then asks whether the pooling method was appropriate; a review that runs a fixed-effect model across obviously heterogeneous studies, or ignores substantial heterogeneity, fails it even if every other item is clean. Because these are critical, a single misstep here is enough to push an otherwise tidy review to low confidence.

Applying AMSTAR 2: a step-by-step appraisal

Treat the appraisal as a reproducible procedure rather than an impression, ideally with two appraisers working independently before they reconcile:

  1. Read the review and its registered protocol side by side, noting where the published methods match or drift from what was planned.
  2. Answer each of the sixteen items as “yes”, “partial yes”, or “no”, citing the exact text or table that supports your answer so the judgement is auditable.
  3. Mark which of your “no” answers fall on critical items, since only these determine the floor of the rating.
  4. Apply the rule: no critical flaws gives high confidence (or moderate if non-critical weaknesses exist), one critical flaw gives low, and more than one gives critically low.

Using AMSTAR 2 alongside GRADE and synthesis

Report the response to every item, not just the overall rating, so a reader can see which domains drove the confidence level. Where AMSTAR 2 judges the review’s conduct, the GRADE certainty rating judges the trustworthiness of its evidence; the two are complementary, and a strong synthesis needs both a credible AMSTAR 2 result and a defensible certainty rating built on the extracted outcome data. This pairing is what an umbrella review of existing reviews relies on: AMSTAR 2 decides which reviews are sound enough to carry forward, and GRADE decides how much weight their conclusions deserve.

Common mistakes when appraising with AMSTAR 2

The most frequent error is treating the overall rating as the whole output. A review reported only as “low confidence” tells a reader nothing about which weakness to worry about; the per-item responses are the substance. A second error is confusing AMSTAR 2 with a study-level bias appraisal: AMSTAR 2 never looks inside the included trials, it judges the method that assembled them. A third is over-rewarding a “partial yes”; partial credit on a critical item is still not a full “yes”, and a string of partials on the search and bias items should leave an honest appraiser cautious rather than reassured. Where you need this done rigorously, whether appraising other teams’ reviews or stress-testing your own before submission, our critical appraisal and risk of bias service works each item against its source text in duplicate.