A systematic review is a literature review that answers a single, clearly defined question by finding every relevant study with a reproducible search, selecting studies against pre-specified eligibility criteria, appraising their risk of bias, and synthesising the results, often with a meta-analysis. What makes it “systematic” is that every decision is planned in a protocol before the evidence is seen, so the conclusions reflect the studies rather than the reviewer’s preferences.

Why the method matters more than the topic

The value of a systematic review comes from its transparency. Two independent teams following the same protocol should reach the same included studies and the same conclusion. That reproducibility is what lets editors, guideline panels, and other researchers trust the result, and it is exactly what a traditional narrative literature review cannot offer. Every stage below exists to remove a specific source of bias, and skipping any one of them weakens the whole review. The discipline is front-loaded: most of the intellectual work happens before you read a single result, when you decide exactly what counts as evidence and how you will handle it.

Before committing, confirm a systematic review is the right design at all. If the literature is large and you want to map what exists rather than answer a narrow effect question, a scoping review may fit better; if you want to pull together existing reviews rather than primary studies, an umbrella review is the right vehicle. The full menu of options, and when each applies, is set out in our guide to the main types of literature review. Choosing the wrong design is the most expensive mistake in the whole process, because it is only obvious months later.

ProtocolSearchScreenExtractAppraiseSynthesisePre-specified in the protocol, then executed in order
The systematic review pipeline: each stage is pre-specified in the protocol before any study is seen.

The stages of a systematic review

1. Frame the question and write a protocol

Everything starts with a precise question. A vague question such as “does exercise help depression” cannot be answered reproducibly, because two reviewers would disagree on which population, which intervention, and which outcome to count. Most reviews therefore structure the question with PICO (Population, Intervention, Comparator, Outcome), and we walk through how to sharpen each element in building a review question. The framework you pick is not cosmetic: it dictates your eligibility rules, your search terms, and the columns of your extraction form.

Once the question is fixed, the methods are written into a formal protocol. The protocol names the databases, the eligibility rules, the primary and secondary outcomes, the risk-of-bias tool, and the planned analysis, so that no decision is made after the results are visible. For health questions the protocol is registered on PROSPERO, which timestamps your methods publicly and lets readers check that your final paper matches what you set out to do. Registration also prevents duplicate effort, because you can see whether another team is already reviewing the same question.

2. Build and run a reproducible search

A systematic review aims to find every eligible study, not a convenient sample, so the search is built to favour comprehensiveness over precision. That means a documented search strategy run across several bibliographic databases such as MEDLINE, Embase, and the Cochrane Library, because no single database indexes everything. Within each, you combine subject headings (controlled vocabulary such as MeSH) with free-text title and abstract terms, joined by Boolean operators: synonyms within a concept linked by OR, separate concepts linked by AND.

A database search alone is rarely enough. You also chase references forwards and backwards through citation searching, and you look for grey literature such as theses, conference abstracts, and trial registry entries to reduce publication bias. Every line of the final search, including dates run and records retrieved, is recorded so the strategy can be rerun and reported. A good test before you run it is independent peer review of the search, which routinely catches a missing synonym or a misplaced operator that would otherwise lose relevant studies.

3. Screen studies in duplicate

The raw search returns far more records than you will keep, and many are duplicates pulled from overlapping databases, so the first step is removing duplicate records. Two reviewers then independently screen titles and abstracts, then full texts, against the eligibility criteria. Working in duplicate catches the studies a single tired reviewer would wrongly discard, which is why most journals expect it and why we explain how many reviewers a review needs. Every full-text exclusion is logged with a reason so the numbers feed cleanly into the PRISMA flow diagram.

Where the two reviewers disagree, a third resolves the conflict, and the level of pre-discussion agreement is quantified with Cohen’s kappa. A low kappa is a useful warning sign that the eligibility criteria are ambiguous and need tightening before screening continues, rather than a problem to paper over at the end.

4. Extract data and assess risk of bias

Reviewers pull the same fields from every included study using a standardised extraction form that has been piloted on a handful of studies first, ideally with two people extracting independently. The form captures study characteristics, participant details, the exact effect estimates, and the variance needed for pooling. When a paper omits a number you need, the answer is often to contact the study authors rather than to guess or drop the study.

In parallel, trained assessors rate each study’s risk of bias with the tool that matches its design: RoB 2 for randomised trials, ROBINS-I for non-randomised studies of interventions, and a tool such as AMSTAR 2 when the included evidence is itself a set of reviews. The judgement is made per domain, not as a single quality score, because only domain-level transparency lets a reader audit why a study was rated high or low.

5. Synthesise the evidence

If the studies are clinically and methodologically similar enough, their results are pooled in a meta-analysis to produce a single weighted estimate, displayed on a forest plot. The consistency of the studies is quantified as heterogeneity using I-squared and the Q statistic; substantial heterogeneity pushes you towards a random-effects model and towards exploring the cause rather than ignoring it. A sensitivity analysis then checks whether the result holds when you remove high-risk studies or change a debatable assumption.

When the studies are too different to pool, forcing a single number is misleading, and the evidence is combined in a structured narrative synthesis instead. Either way, the overall certainty of the evidence is rated with GRADE, which downgrades for risk of bias, inconsistency, indirectness, imprecision, and publication bias so that the strength of your conclusion is stated explicitly rather than implied.

6. Report against PRISMA

The finished review is written up against the PRISMA 2020 checklist, with the flow diagram, the full search for at least one database, the risk-of-bias judgements, and the synthesis all laid out so a reader can audit every step. PRISMA is a reporting standard, not a quality scale: it does not tell you whether the review was done well, but it forces you to disclose enough that others can judge for themselves. The fastest route to a clean write-up is to keep the protocol, the search log, and the screening records as you go, rather than reconstructing them at the end.

What to settle before you write a single search line

Most of the time lost on a review is spent reworking decisions that should have been fixed at the start. Before you touch a database, settle the following, and write each into the protocol so it cannot quietly change later:

  • The exact question in PICO terms, including which single primary outcome the review is powered to answer.
  • The inclusion and exclusion criteria, down to study designs, dates, languages, and settings.
  • Which databases and supplementary sources you will search, and who will design the strategy.
  • The risk-of-bias tool matched to your study designs, chosen before you know the results.
  • Whether you expect to pool the data, and the model and effect measure you will use if you do.
  • Who fills each role: two screeners, two extractors, a tie-breaker, and ideally an information specialist and a statistician.

A team that can answer all six before starting will move through the rest of the project with far less friction, because every later stage simply executes a decision that has already been made.

Writing up the manuscript, section by section

A systematic review manuscript follows a predictable shape, and knowing what each section must contain prevents the scramble at submission. The abstract is usually structured and should state the question, the eligibility criteria, the number of studies and participants, the main effect estimate with its confidence interval, and the certainty rating. The introduction justifies why the review is needed now and ends on the precise question, ideally in PICO terms, so the reader knows exactly what will be answered.

The methods section is where reviewers spend most of their scrutiny: it reports the registration number, the full eligibility rules, every database and the dates searched, the screening process, the risk-of-bias tool, and the planned synthesis, all matching the protocol. The results section opens with the study-selection flow, then describes the included studies, their risk of bias, and the synthesis, with the forest plot and any subgroup or sensitivity analyses. The discussion interprets the pooled estimate in light of its certainty, states the limitations honestly, and avoids overclaiming. A disciplined manuscript reads as the protocol carried through to its conclusion, with no surprises that were not pre-specified.

Assembling the team and keeping the review reproducible

Although a single researcher can lead a review, the method assumes more than one pair of hands. At minimum you need two screeners and two extractors, with a third person to break ties, and ideally an information specialist to design the search and a statistician for any pooling. The review will be stronger and faster when those roles are assigned at the protocol stage rather than improvised mid-project, because each stage hands a clean, documented output to the next.

Reproducibility is not a single step but a habit maintained throughout. Keep the exact search strings, the export of records at each screening stage, the reasons for every full-text exclusion, and a dated log of any protocol amendment. If a reviewer, an editor, or a guideline panel later asks how you reached a particular included set, you should be able to answer from your files in minutes. That same archive is what makes a review straightforward to update in two or three years, when new trials have appeared and the question is worth revisiting.

Common mistakes that get a review rejected

Most rejections trace back to a small number of avoidable errors. The first is a question that drifts: the final paper analyses something the protocol never specified, which reviewers spot immediately by comparing the manuscript to the registration. The second is a thin search run in a single database, which cannot credibly claim to have found all the evidence. The third is single-reviewer screening, which introduces exactly the selection bias the method exists to prevent.

Two more are statistical. Pooling studies that should never have been combined produces a precise-looking estimate that means nothing, and treating a high risk-of-bias study as if it were sound lets a flawed trial dominate the result. A review that registers its protocol, searches broadly, screens in duplicate, appraises bias honestly, and pools only when pooling is defensible will clear peer review on method even when the underlying evidence is weak, because weakness in the evidence is a finding rather than a fault.

How long it takes and how to make it manageable

A full systematic review is a substantial project. A realistic timeline runs from a few months to well over a year depending on the breadth of the question and the volume of records, as we cover in how long a systematic review takes. The two reliable ways to keep it manageable are a tight, well-framed question that limits the record count, and dividing the labour-intensive stages, screening and extraction, across more than one reviewer so the work runs in parallel rather than queueing behind one person.

If you want the shorter version of the lifecycle as a checklist, see the steps in a systematic review. And if any single stage, or the whole project, is more than your team can carry alongside its other work, that is exactly the kind of support we provide to a registered, fully reported standard.