Inclusion and exclusion criteria are the pre-specified rules that decide which studies enter your systematic review and which are kept out. They define the eligible populations, study designs, comparators, outcomes, languages, and dates, and because they are fixed in the protocol before screening, they keep selection objective rather than convenient.

Why criteria are the boundary that defines your review

Your eligibility criteria are, in effect, the operational definition of your question. If the criteria are loose, screening becomes a series of judgement calls that two reviewers will make differently; if they are tight and explicit, screening becomes consistent and auditable. They flow directly out of your PICO question, with each PICO block becoming a criterion, and they feed straight into title and abstract screening. Set them well in the protocol and the rest of selection runs almost mechanically.

All recordsCriteriaappliedIncludedExcluded
Eligibility criteria are the filter that sorts every record into included or excluded, with a recorded reason.

The dimensions every set of criteria should cover

Population, intervention, and comparator

Start from your question. State who counts as eligible (age, condition, setting), what intervention qualifies, and what comparators are acceptable. A common mistake is being vague about the population, which then lets borderline studies slip in and inflates your heterogeneity later.

Study design

Name the eligible designs explicitly: randomised trials only, or observational studies too, and so on. This single decision shapes which risk of bias tool you will use, since RoB 2 and ROBINS-I apply to different designs. Be explicit about whether conference abstracts and grey literature count.

Outcomes, language, and dates

Specify which outcomes a study must report to be included, and decide on any language or date limits. Language limits are sometimes necessary but should be justified, because excluding non-English studies can introduce bias. Date limits should track a real event, such as a change in practice, not just convenience.

Why the population is the criterion most often botched

Of all the dimensions, the population is where loose wording does the most damage, because it is the criterion reviewers feel they understand intuitively and therefore leave underspecified. A rule such as “patients with depression” hides a dozen unstated decisions: does it include subclinical symptoms or only a formal diagnosis, treatment-resistant cases or first episodes, inpatients or community samples? Each unanswered question becomes a point where two reviewers diverge. A population that names the diagnostic threshold, the care setting, and any age band converts those silent judgement calls into a single shared rule, which is the whole purpose of writing the criteria down before screening. The same precision keeps the included studies comparable enough that any later numeric synthesis is defensible rather than a forced average across mismatched groups.

Inclusion versus exclusion: avoid the double count

A frequent error is listing the same rule on both sides, for example “adults” as an inclusion criterion and “children” as an exclusion criterion. That redundancy clutters screening and invites contradiction. As a rule, define each dimension once as an inclusion criterion, and reserve true exclusion criteria for cases that would otherwise pass, such as wrong study design or duplicate publication of the same dataset. Recording the reason for every full-text exclusion is what populates your PRISMA flow diagram.

A worked set of criteria

Criteria read most clearly as a paired list, each inclusion rule shadowed by the exclusion it implies. For a review of statins for primary prevention, a defensible set looks like this:

  • Population: adults with no prior cardiovascular event. Excluded: secondary-prevention populations, because they answer a different question.
  • Intervention: any statin at a stated dose. Excluded: studies combining a statin with another lipid-lowering drug, where the effect cannot be attributed.
  • Comparator: placebo or no treatment. Excluded: head-to-head trials of two active drugs, which lack the contrast the question needs.
  • Study design: randomised controlled trials only. Excluded: observational designs, because the appraisal tool and confounding risks differ.
  • Outcome: must report a cardiovascular event or mortality. Excluded: studies reporting only a surrogate marker such as cholesterol level.

Notice that every exclusion is the necessary mirror of an inclusion rule, not a free-floating extra. That discipline is what stops the redundant double counting described above and keeps the list short enough for two reviewers to hold in their heads during a fast first pass.

Common mistakes that fracture screening

Most disagreements between reviewers trace back to a small set of avoidable flaws in the criteria themselves:

  1. Unstated thresholds. “Older adults” or “severe disease” without a number forces each reviewer to invent a cut-off, so two people screen against two different rules.
  2. Outcome rules hidden in the population block. Mixing a required outcome into the population definition lets studies pass title screening that should never have entered, inflating the full-text workload.
  3. Convenience language limits. Restricting to English for ease rather than for a justified reason can introduce selective inclusion and is flagged by reviewers as a limitation.
  4. Criteria that drift mid-screen. Quietly reinterpreting a rule once results start appearing is exactly the bias the protocol exists to prevent. Any genuine change must be dated and recorded.

A set that survives these tests turns selection into something close to mechanical, which is the standard our protocol service writes toward so that two independent reviewers reach the same call without negotiation.

Piloting before you commit

Criteria that look clear on paper often fracture on real abstracts. Run a small pilot: have two reviewers apply the draft criteria to the same handful of records and compare. Disagreements reveal ambiguous wording you can fix before the full screen. Measuring early agreement with inter-rater reliability tells you whether the criteria are genuinely shared, and a quick kappa calculation on the pilot records puts a number on it. If you refine them after registration, record the change transparently, as we cover in amending a protocol.