Piloting a data extraction form means testing your draft form on a small sample of included studies before full extraction begins, so that ambiguous fields, missing categories, and inconsistent interpretations are found and fixed while they are still cheap to correct. It is a dress rehearsal for the real extraction, and skipping it almost always means re-doing work once the same problem surfaces across dozens of studies.

What a pilot is actually testing

A pilot is not a formality. It checks three different things at once: whether the form captures everything your synthesis needs, whether two reviewers reading the same paper record the same answers, and whether the field definitions hold up against the messy reality of published reports. Catching a flaw here, before it contaminates the full data extraction run, protects every result that follows, from the risk of bias judgements to the final meta-analysis.

Extract sampleCompareRevise formFull runLoop until reviewers agree, then commit
A pilot loops: extract a small sample, compare, revise the form, and only then commit to full extraction.

How to run the pilot, step by step

Choose a representative sample

Pick a handful of included studies that span your evidence base, different designs, different reporting styles, and at least one awkward paper. A pilot run only on the cleanest studies will pass easily and teach you nothing. Five to ten studies is usually enough to expose the recurring problems.

Have two reviewers extract independently

Each reviewer fills the form on the same pilot studies without conferring, exactly as they will during the real run. This is the same logic that underpins duplicate full-text screening: silent disagreement only shows up when two people work in parallel. The point is to see where the form lets them diverge.

Compare entries and measure agreement

Lay the two sets of answers side by side, field by field. Where they disagree, ask whether the fault is the reviewer or the form. For categorical fields you can quantify consistency with a Cohen’s kappa calculator, the same way you would gauge inter-rater reliability in screening. Low agreement on a field is a signal that the field, not the reviewers, needs rewriting.

Acting on what the pilot reveals

Fix ambiguous fields and definitions

Most pilot failures trace back to a vague field. Rewrite the definition, split a field that was trying to do two jobs, or add a category you had not anticipated. Update the codebook so the new rule is written down, and keep the form aligned with the eligibility criteria from your protocol.

Calibrate the reviewers

Some disagreements are not about the form but about how reviewers apply it. Walk through the discrepancies together, agree the correct reading, and note the resolution, exactly as you would when resolving screening conflicts. A short calibration discussion now prevents systematic drift later.

Decide whether a second pilot is needed

If the first pilot forced major changes, run a quick second round on fresh studies to confirm the revisions worked. If only minor wording changed, you can usually proceed. Either way, record that the form was piloted, since reporting it is part of a transparent systematic review process and tells reviewers your dataset rests on a tested instrument.

A practical pilot checklist

A pilot is more reliable when it follows a fixed routine rather than a general intention to “try the form out”. Work through these steps in order and the pilot will surface the problems that matter:

  1. Select five to ten included studies that span your designs, with at least one poorly reported paper deliberately included.
  2. Give both reviewers the current form, the codebook, and the same set of full-text articles, and have them extract independently with no conferring.
  3. Merge the two sets into a single comparison sheet, one row per field per study, so every disagreement is visible at a glance.
  4. Classify each disagreement as a form fault (ambiguous definition, missing category, wrong format) or a reviewer fault (a genuine misreading), because the two need different fixes.
  5. Time how long one study takes to extract, then multiply by your included count to sanity-check whether the form is realistic for the full set.
  6. Revise the form and codebook, log every change with a reason, and decide whether a second pilot round is warranted.

What the disagreements are telling you

The pattern of mismatches is diagnostic. A field where reviewers split roughly evenly is genuinely ambiguous and needs its definition rewritten or its categories split. A field where one reviewer is consistently high or low points to systematic drift in how one person reads the instruction, which a calibration discussion fixes. A field that is blank for one reviewer but filled for the other usually means the data is hard to find in the paper, which is a signal to add an explicit location prompt to the codebook, for example “extract from Table 2, not the abstract”.

Quantifying this matters most for the categorical judgement fields. A Cohen’s kappa below about 0.6 on a field is a clear instruction to revise rather than proceed, the same threshold logic you apply to dual-reviewer screening decisions. Continuous fields, such as a mean or a sample size, are better checked by the raw count of exact mismatches, since a single transposed digit will not move a kappa but will corrupt a pooled estimate.

How a pilot saves time on the full run

It can feel like a detour to extract five studies twice before the real work starts, but the arithmetic favours piloting. A single ambiguous field that slips through to a full run of sixty studies must be re-read across all sixty, doubling back through papers reviewers have already closed, often weeks later when the original reasoning has faded. The same fault caught in a five-study pilot costs one short rewrite. The pilot also calibrates pace: by timing how long one study takes, you can forecast the full extraction and decide whether to recruit a second pair of extractors before the queue backs up. That forecast feeds directly into the realistic scheduling we describe in how long a review takes.

A pilot is also the moment to confirm the form actually produces the inputs your analysis expects. If your plan is to pool continuous outcomes, the pilot should yield a clean mean, standard deviation, and sample size for every arm, so that a missing-variance problem surfaces now rather than when you load the dataset into a forest plot. Treating the pilot output as a dry run of the synthesis, not just a test of the form, is what makes the rest of the project flow.

Common piloting mistakes

  • Piloting only on clean, well-reported trials, so the form passes and then collapses on the first messy observational study.
  • Letting the two reviewers confer during the pilot, which hides the very disagreements the exercise is meant to expose.
  • Fixing the wording in reviewers’ heads but not in the written codebook, so the rule is lost the moment a new reviewer joins.
  • Skipping the second pilot after major revisions and assuming the fix worked without testing it.
  • Treating the pilot as a one-off when the form draws on tools that change how data is captured, such as one of the review management platforms with built-in extraction modules.

When the pilot exposes deeper problems

Occasionally a pilot reveals that the data you need simply is not reported in a usable form, only inside charts, or missing entirely. That is your cue to plan for extracting data from figures or for contacting authors for missing data before full extraction, so the workaround is built into your method rather than improvised mid-stream. A pilot can even surface a problem upstream of the form: if reviewers keep disagreeing about whether a study belongs at all, the real fault may lie in vague eligibility rules rather than the extraction fields, and that is far cheaper to correct now than after the dataset is half built.