The Newcastle-Ottawa Scale is a star-based instrument for appraising the quality of observational studies, chiefly cohort and case-control designs, in a systematic review. It awards stars across three groups, the selection of study groups, the comparability of those groups, and the ascertainment of the outcome or exposure, giving each study a simple score that flags how vulnerable it is to bias.
Where the scale fits among appraisal tools
Observational studies cannot be appraised with the RoB 2 tool, which assumes randomisation. Some teams reach for ROBINS-I, but its target-trial logic is heavy for descriptive cohort and case-control questions. The Newcastle-Ottawa Scale offers a lighter, widely accepted alternative that still produces a defensible judgement, which is why it appears in so many reviews of non-interventional evidence as part of the wider risk of bias assessment.
The three domains and their stars
Selection
Up to four stars reward how well the study chose its groups: whether the exposed cohort was representative, how the unexposed group was drawn, how exposure was ascertained, and, for cohorts, whether the outcome was absent at the start. Weak selection here is the observational equivalent of broken randomisation.
Comparability
Up to two stars are given for how the study controlled confounding, one for adjusting for the single most important factor and a second for adjusting for others. This is the single domain where reviewers most often disagree, because the scale leaves the review team to name which confounders count, ideally in the protocol.
Outcome or exposure
The final group awards up to three stars for how the outcome (in cohort studies) or exposure (in case-control studies) was determined, the length and adequacy of follow-up, and how non-response was handled. Outcomes confirmed by records or independent assessment score better than those based on self-report.
A worked scoring example
Picture a cohort study comparing an exposed and an unexposed group. It draws the exposed cohort from a truly representative community sample (one star), draws the unexposed group from the same community (one star), confirms exposure from secure medical records rather than self-report (one star), and shows the outcome was absent at baseline (one star), giving a full four stars for selection. It adjusts for age and sex, the single most important confounder set for its question (one star), but not for smoking or comorbidity (no second star), so it earns one of two comparability stars. It confirms the outcome by linked records (one star), follows participants long enough for the outcome to occur (one star), but loses more than a fifth of the cohort to follow-up with no description of the dropouts (no star), earning two of three outcome stars. The total is seven of nine. Reported as a bare “seven” that looks strong, yet the missing follow-up star is exactly the attrition weakness a reader most needs to see, which is the case for reporting the stars by domain.
Interpreting the score, and its limits
Many reviews convert the star total into good, fair, or poor bands. That convenience is also the scale’s main weakness: a single number can hide a serious flaw in one domain behind strength in another, and the bands are not standardised across reviews. Best practice is to report the stars per domain, not just the total, and to define your own thresholds in advance so the scoring cannot drift to suit the conclusion. This is also where the difference between true systematic error and broader quality features becomes practical rather than academic.
A second limitation is the loose anchoring of the comparability stars. Because the scale leaves the team to name which confounders earn a star, two reviews of the same studies can award different stars unless the criteria are fixed first. Pin them down in the review protocol (for example, “one star for age and sex, a second for a measure of baseline disease severity”) and have two assessors score independently before reconciling. Where pre-discussion agreement is low, quantify it with a chance-corrected agreement statistic and tighten the rules before continuing. The same discipline that governs full-text screening decisions applies to appraisal: a judgement that cannot be reproduced is not a judgement.
How the scale compares with newer instruments
The Newcastle-Ottawa Scale predates the design-matched tools and remains popular for its speed, but it is worth knowing what it trades away. Unlike the target-trial approach of ROBINS-I, it does not force the team to specify a causal contrast or a full confounder set, which is why it is lighter but also blunter on confounding. Unlike the signalling-question structure of RoB 2, it has no built-in algorithm, so its judgements lean more on the assessor. For a descriptive review of cohort and case-control evidence it is a defensible choice; for a review whose conclusions hinge on a causal claim from observational data, many teams now prefer the heavier tool and accept the extra work. Whichever you pick should be named in advance as part of the wider appraisal plan for the review, never swapped after the results are in view.
The case-control version and its differences
The Newcastle-Ottawa Scale comes in two forms, and reaching for the wrong one is a quiet but real error. In the cohort version, the selection stars reward how the exposed and unexposed groups were drawn and whether the outcome was absent at baseline, and the final group judges the outcome and the adequacy of follow-up. In the case-control version, the logic flips: cases and controls are defined by their outcome, so the selection stars reward an adequate case definition and a sensibly chosen control group, and the final group judges how exposure, rather than outcome, was ascertained, including whether the same method was used for cases and controls. Mixing the two, for example asking about follow-up adequacy in a case-control study where there is no follow-up, produces stars that do not mean what the table claims. Choose the version that matches each included design and say which you used, just as you would match the instrument to the design across the rest of the review.
Carrying the judgement forward
The per-domain stars feed the study-limitations criterion in the GRADE certainty rating, and the ratings are reported beside the extracted study data so a reader can see exactly why an observational result was weighted as it was in the narrative or pooled synthesis. When several studies share the same domain weakness, that pattern, rather than any individual total, is what drives a downgrade for study limitations, and it is the reason a transparent per-domain table is worth far more than a column of nine-point scores. If applying the scale consistently across a body of observational studies is more than your team can carry, our observational study appraisal service pre-defines the criteria and scores in duplicate so the stars hold up at review.