Data extraction in a systematic review is the process of pulling the same pre-specified fields from every included study into a structured form, so that study characteristics, methods, and outcome results can be compared and synthesised. It turns a stack of full-text papers into a clean, analysable dataset, and because every field is fixed in advance, two reviewers working independently should record the same numbers from the same paper.

Why extraction is the load-bearing stage of the review

Every downstream result depends on what comes out of this stage. A meta-analysis can only pool the effect estimates that were extracted correctly, and a narrative synthesis is only as honest as the table behind it. If a sample size is mistyped or an outcome is read from the wrong column, that error propagates silently into the forest plot and the conclusion. This is why structured data extraction is treated as a method to be planned, piloted, and audited, not a clerical afterthought.

ScreenExtractReconcileSynthesiseEach field is fixed in advance and extracted in duplicate
Extraction sits between screening and synthesis: structured fields flow from each study into one shared dataset.

What to extract, and how to organise the fields

Study and population characteristics

Begin with the descriptive backbone: author, year, country, study design, setting, funding source, and the sample characteristics that define your PICO elements. These fields let readers judge applicability and let you build the characteristics of included studies table that every systematic review needs. Capturing design here also feeds directly into your later risk of bias assessment, since the right appraisal tool depends on the design.

Methods and outcome definitions

Record how each outcome was measured, at what time points, and with what instrument. Two studies may both report “anxiety,” but if one uses a different scale or follow-up window, you need that detail to decide whether they are poolable. Aligning these definitions against the eligibility rules you set in your inclusion and exclusion criteria keeps the dataset consistent.

Numerical results for synthesis

For each outcome, extract the raw numbers your synthesis needs: events and totals for binary outcomes, or means, standard deviations, and sample sizes for continuous ones. Where a paper reports a standard error, confidence interval, or p-value instead, note exactly what was given so it can be converted later with an effect size converter rather than guessed at.

Extracting in duplicate and reconciling

The reliable approach is two reviewers extracting each study independently, then comparing entries field by field and resolving any mismatch by discussion or a third reviewer. This mirrors the logic of duplicate full-text screening: a single reader makes silent slips, and a second pass catches them. Recording who extracted what, and how conflicts were resolved, gives your review the same auditability you build during conflict resolution at screening.

How much to extract, and what to leave out

A recurring tension is breadth versus discipline. Extract too little and you return to the papers when a reviewer asks a question you did not anticipate; extract too much and reviewers tire and make errors on the fields that actually matter. The resolution is to tie every field to a planned use. If a field will not appear in the characteristics table, feed an appraisal judgement, or enter the synthesis, it probably does not belong on the form. Map each field back to the elements of your review question and to the certainty-of-evidence domains you will rate, and the list trims itself.

Be especially deliberate about the variance fields. The most common reason a study cannot be pooled is that the measure of spread was never captured, so build explicit slots for the standard deviation or for whatever the paper reported instead, and flag the format so it can be converted with an odds ratio and risk ratio calculator or a comparable derivation rather than estimated by eye.

Single versus duplicate extraction

Reviewers sometimes ask whether full duplicate extraction is always necessary. The gold standard is two independent extractors on every field of every study, because a single reader makes silent transcription slips that nobody catches. Where resource is tight, a defensible middle ground is one person extracting and a second verifying against the source paper, which catches most errors while halving the second-reader effort, though it is weaker than true independence. What is never acceptable is a single unchecked extractor on the numbers that drive the result, which reintroduces exactly the error the method exists to remove. We run full duplicate extraction as part of our data extraction service for precisely this reason.

Common extraction errors and how to catch them

  • Unit and scale slips: recording a value in different units across studies, or reading a percentage as a proportion. A units field per outcome forces the question at the point of entry.
  • Wrong arm or wrong comparator: extracting the control arm as the intervention. A codebook rule naming which arm is which prevents it.
  • Confusing spread measures: entering a standard error where a standard deviation is expected, which silently inflates a study’s weight in the synthesis.
  • Double-counting linked reports: treating two papers from one trial as two studies. Collapse them into a single record before synthesis.
  • Intention-to-treat versus per-protocol: mixing analysis populations across studies, which a dedicated field makes visible.

Building and testing the form

A good form is the difference between clean and chaotic extraction. Start from a reusable data extraction form template, tailor the fields to your question, and always run a pilot on a handful of studies before full extraction. Piloting surfaces ambiguous fields and missing categories while they are still cheap to fix. When numbers appear only inside charts, you may need techniques for extracting data from figures, and when a paper omits the statistic you need, the next step is often contacting the authors.

Where extraction fits in the wider workflow

Extraction is one stage in a planned sequence. It follows screening and feeds appraisal and synthesis, all of it laid out in advance, as we cover in the steps of a systematic review. Treating the dataset as a deliverable in its own right, complete with a codebook and an audit trail, is what lets a guideline panel or journal editor trust the numbers that follow.