Extracting data from figures is the process of recovering numerical values from a published chart, such as a bar chart, line graph, scatter plot, or Kaplan-Meier curve, when the underlying numbers are not reported in the text or tables. Using a calibrated digitising tool, you map plotted points back to the axes to read off the means, proportions, or survival estimates your synthesis needs, then record them like any other extracted field.

Why reviewers end up reading graphs

Plenty of studies present a result only as a figure. The mean at each time point lives in a line graph, the event rate sits in a bar chart, and no table gives you the number. Excluding those studies would bias your review toward papers that happened to tabulate their results, so recovering the values keeps your data extraction complete. The recovered numbers then flow into the same pipeline as everything else, feeding the meta-analysis or the narrative summary.

recovered valueCalibrate axes, then read each point
Digitising maps a plotted point back to the calibrated axes to recover its underlying value.

How to digitise a figure reliably

Calibrate the axes first

Open the figure in a digitising tool such as WebPlotDigitizer and mark known reference points on each axis, two points of known value on the x-axis and two on the y-axis. This calibration is what lets the software convert pixel positions into real units. A log-scaled axis needs to be set as logarithmic during calibration, or every value you read will be wrong.

Read the data points

With the axes calibrated, click each data point and let the tool report its coordinates. For a line graph, capture the value at every reported time point; for a bar chart, read the top of each bar; for a scatter plot, you may capture every point or a summary. Where error bars are shown, read them too, since they often encode the standard deviation or interval you need for synthesis.

Convert what you recover into usable statistics

A recovered mean is only half the picture. If the figure shows a confidence interval rather than a standard deviation, reconstruct the standard deviation with a confidence interval calculator, and if you have a different effect metric than the one your model needs, translate it with an effect size converter. The aim is to land on the same fields your extraction form expects for every other study.

Recovering numbers from common figure types

Each chart type hides its numbers differently, and the recovery routine changes accordingly. Knowing the trap each one sets is what separates a defensible reading from a guess.

  • Bar charts. Read the top edge of each bar against the y-axis. Watch for a y-axis that does not start at zero, which exaggerates differences, and for stacked bars, where you must subtract one segment from the cumulative total to recover the individual value.
  • Line graphs. Capture the value at every reported time point, not just the endpoints, because a synthesis at one follow-up window needs that exact point. Where two lines overlap, zoom in before clicking so you assign the point to the correct series.
  • Box plots. Recover the median, the upper and lower quartiles, and the whiskers, then convert that five-number summary into a mean and standard deviation using a published estimator before pooling, since most models expect the mean.
  • Scatter plots. Digitise every point when you need the raw distribution, for example to recompute a correlation, and record the count so a reader can confirm the sample size matches the text.
  • Kaplan-Meier curves. These need the most care. Digitise the step function and combine it with the numbers-at-risk table printed below the plot to reconstruct individual time-to-event data, following the Guyot method, rather than reading a single survival probability off the curve.

A worked mini-example

Suppose a trial plots mean pain score over twelve weeks with the y-axis running 0 to 100 and gridlines every 20. You calibrate by marking the 0 and 100 points on the y-axis and weeks 0 and 12 on the x-axis. Clicking the treatment line at week 12 returns 34, and the error bar caps sit at 28 and 40. Because the caption says the bars are 95 per cent confidence intervals on a sample of 50, you convert the half-width of 6 into a standard deviation: the standard error is 6 divided by 1.96, giving about 3.06, and multiplying by the square root of 50 returns a standard deviation of roughly 21.6. That mean and standard deviation now slot straight into the continuous-outcome slots of your structured extraction dataset.

Common digitising mistakes that corrupt a result

A handful of errors account for most bad readings, and every one of them is avoidable with a slower, more deliberate routine:

  • Leaving an axis on a linear setting when the figure is log-scaled, so an odds ratio of 2 is recorded as if the spacing were arithmetic.
  • Confusing a standard error bar with a standard deviation bar or a confidence interval; the caption, not the appearance, tells you which, and the conversion differs by the square root of the sample size.
  • Calibrating on a low-resolution screenshot rather than the publisher PDF, where a few blurred pixels at the axis become several units of error.
  • Reading a cumulative or percentage axis as if it were a count, or missing a break in the axis that compresses part of the range.

When you correctly recover the variance alongside the point estimate, those studies can carry their proper weight in a forest plot instead of being down-weighted by an inflated standard error or, worse, dropped altogether.

Keeping digitised data trustworthy

Digitise in duplicate

Reading a graph by eye introduces error, so have two reviewers digitise the same figure independently and compare. If their readings agree closely, take the average; if they diverge, investigate before trusting either. This is the same duplicate-and-reconcile discipline you apply across full-text screening and the rest of extraction.

Document the method and its limits

Record which tool you used, how you calibrated, and which values were digitised rather than reported. Digitised data carries more uncertainty than tabulated data, which matters when you weigh it in a meta-analysis calculator and when you later run a sensitivity analysis to check whether those studies sway the result. Flagging the source also keeps your risk of bias reporting honest.

When digitising is not enough

Sometimes a figure is too small, too cluttered, or simply does not contain the value you need. Rather than guess, treat it as missing data and move to contacting the authors. Requesting the raw numbers is more defensible than a strained reading, and it keeps your systematic review process transparent about where each figure came from.