Data extraction and coding services

We pilot the form, extract in duplicate, recover numbers from figures, and chase authors for missing data, so the dataset behind your synthesis is accurate and defensible.

4.9 / 5 across 1,194+ delivered projects

  • Cochrane and PRISMA 2020 methods
  • PhD methodologists
  • 100% human-written, no generative AI
  • Reproducible R and Stata code
  • Free quote within 24 hours
  • Mutual NDA on request

Data extraction is the stage where the information needed for your synthesis is pulled from each included study into a structured dataset. Our service builds and pilots the extraction form, runs duplicate independent extraction with reconciliation, recovers values from figures when text is missing, contacts authors for unreported data, and applies consistent coding so every study enters your analysis on the same terms.

Why a clean extraction is the foundation your results sit on

A meta-analysis is only as trustworthy as the numbers fed into it. A single mistyped standard deviation, a misread time point, or an outcome coded inconsistently across studies can move a pooled estimate and survive all the way to publication unnoticed. This is the stage where errors are cheapest to catch and most expensive to ignore, which is why it is judged on process: piloting, independence, and documented reconciliation rather than one person filling in a spreadsheet. We build that discipline in, starting from a structured extraction form template tailored to your review.

Piloting matters more than it looks. Running the draft form on a handful of studies exposes ambiguous fields and outcome definitions before they contaminate the full dataset. Our approach to piloting a data extraction form fixes those problems early, and where your literature search and screening have already produced the included set, extraction follows directly from it.

What the extraction service covers

  1. 1

    Form design and piloting

    We design an extraction form around your outcomes and risk of bias needs, then pilot it on a sample of studies and refine the fields before full extraction.

  2. 2

    Duplicate independent extraction

    Two reviewers extract each study independently, then reconcile differences against the protocol, with every disagreement logged and resolved.

  3. 3

    Recovering and deriving data

    We digitise values from figures, derive missing statistics from what is reported, and flag every estimated number in the dataset.

  4. 4

    Author contact and coding

    We contact authors for unreported data, track each request, and apply consistent coding so outcomes and subgroups align across studies.

The tools and techniques behind a clean dataset

Review management platforms

We extract in Covidence, EPPI-Reviewer, or a structured spreadsheet, matching whatever your team already uses for the review.

Calibrated figure digitising

When an outcome appears only in a graph, we recover the values with calibrated digitising and flag every estimated number.

Statistic derivation

We compute missing standard deviations from confidence intervals, standard errors, or p-values where the reporting allows.

Agreement measurement

We quantify reviewer agreement on screening and extraction so the dual-reviewer process is documented, not just asserted.

What you receive

You receive a clean, fully documented extraction dataset, the piloted form, a reconciliation log showing how disagreements were resolved, and a record of every author contact and figure-derived value. The dataset is structured to flow straight into risk of bias and quality assessment and into the statistical analysis and synthesis. If you also need the reviewer-agreement statistics for screening and extraction, our free Cohen’s kappa calculator and our guide to inter-rater reliability in screening show how to report agreement properly.

  • A clean extraction dataset ready for synthesis
  • The piloted, review-specific extraction form
  • A reconciliation log recording how each disagreement was resolved
  • Reviewer-agreement statistics for the extraction process
  • A documented trail of every figure-derived value and author request
  • Consistent coding so outcomes and subgroups align across studies

Who needs extraction handled for them

Solo reviewers needing a second extractor

Researchers who cannot run a credible dual-reviewer process alone and need an independent second extraction.

Teams under a deadline

Groups with a large included set who need accurate extraction completed without stalling the rest of the review.

Doctoral candidates

Students who need a defensible, documented dataset that will hold up when examiners probe how the numbers were obtained.

Authors with messy reporting

Reviews where studies report outcomes in graphs or incompatible formats that need recovering and harmonising.

How extraction work is quoted

We quote a fixed fee based on the work the dataset requires: how many studies are included, how many outcomes and time points each reports, whether values must be recovered from figures or derived from partial statistics, and whether author contact is likely. Send your included studies and your outcome list and we will return a fixed quote for piloted, duplicate extraction.

Single extraction yourself versus duplicate extraction with us

ConsiderationSingle extractionDuplicate extraction with us
Error rateTranscription and judgement errors go undetectedTwo independent passes catch mistakes one reviewer would miss
DefensibilityHard to evidence how a contested number was obtainedA reconciliation log and agreement statistics back every field
Missing dataOften left blank or guessed atDerived where possible, requested from authors where not, and flagged
Downstream analysisA flawed dataset can move a pooled estimate without warningA clean dataset that flows straight into synthesis

Frequently asked questions

Why does data extraction need two independent reviewers?
Single extraction introduces transcription and judgement errors that change a pooled result. Two reviewers extracting independently and then reconciling differences catches mistakes one reviewer would miss and produces a documented, defensible dataset. We run duplicate extraction by default and record how every disagreement was resolved.
What should be on a data extraction form?
An extraction form captures the study identifiers, the population and setting, the intervention and comparator, the outcomes and time points, the effect estimates or raw numbers, the sample sizes, and the items needed for risk of bias assessment. We design the form around your specific review and pilot it on a sample of studies before full extraction begins.
Can you extract data that is only shown in a figure?
Yes. When a study reports an outcome only in a graph, we use calibrated digitising methods to recover the values, and we note in the dataset that the number was estimated from a figure. Where the figure cannot be read reliably, we contact the authors for the underlying data.
What happens when a study does not report the data you need?
We first try to derive it from what is reported, for example calculating a missing standard deviation from a confidence interval or a standard error. When that is not possible, we contact the study authors and request the missing values, and we document every request and response so the process is transparent.

Send us your included studies and your outcomes, and we will pilot a form and extract your data in duplicate for a fixed quote.

Free quote within 24 hours. 100% human-written by PhD methodologists.

Methodology reviewed by

Dr Eleanor Whitfield, PhD

Lead Review Methodologist

Leads protocol design, screening, and reporting across the practice, with a focus on reviews that survive editorial scrutiny.

Why researchers bring in a PhD methodologist

80+

systematic reviews are published every day (Hoffmann et al., 2021)

67.3 weeks

average time to complete a review in-house (Borah et al., BMJ Open 2017)

~70%

of published reviews rate critically low on AMSTAR 2 quality appraisal

4.9 / 5

our client rating across 1,194+ delivered projects