To do a meta-analysis you extract a comparable effect size from every included study, weight each study by the precision of its estimate, and combine them into a single pooled estimate with a confidence interval. The weighting means larger, more precise studies count more, and the result is reported with a measure of heterogeneity that tells you how much the true effects vary across studies. A meta-analysis sits inside a systematic review; it is the quantitative synthesis stage, not a separate project.
When pooling is justified and when it is not
The first decision is whether the studies are similar enough to combine at all. If the populations, interventions, and outcomes are clinically and methodologically comparable, a quantitative synthesis is appropriate. If they are too different, forcing a single number hides more than it reveals, and a structured narrative synthesis is the honest choice. This judgement is made before you see the pooled result, and it is the line that separates a defensible systematic review with a meta-analysis from a misleading one.
Choosing a common effect measure
Every study has to report on the same scale before anything can be pooled. For binary outcomes the usual measures are the odds ratio, the risk ratio, or the risk difference; the trade-offs between the first two are covered in our guide to odds ratio versus risk ratio. For continuous outcomes you use a mean difference when every study measured the outcome on the same instrument, or a standardised mean difference when instruments differ. The full menu of options is set out in effect sizes in meta-analysis. Once the measure is fixed, you can run the numbers quickly with our meta-analysis calculator.
Weighting and the inverse-variance principle
The engine of any meta-analysis is the weight each study receives. The standard approach is inverse-variance weighting: a study with a small standard error, meaning a precise estimate, gets a large weight, while an imprecise small study gets a small one. This is why two well-powered trials can dominate a pool of a dozen tiny ones. The pooled estimate is simply the weighted average of the individual effects, and its confidence interval narrows as the combined sample size grows.
A small worked example makes the mechanics concrete. Imagine three studies of the same continuous outcome, reporting effects of 0.30, 0.50, and 0.45 with standard errors of 0.20, 0.10, and 0.15. Their inverse-variance weights are one divided by each squared standard error, giving 25, 100, and roughly 44. The precise middle study carries more than half the total weight. Multiplying each effect by its weight, summing, and dividing by the total weight yields a pooled effect near 0.46, pulled firmly toward the most precise study rather than sitting at the simple average of 0.42. The pooled standard error is the square root of one divided by the summed weights, which is why adding precise studies tightens the interval. For binary outcomes the same logic runs on the log scale for odds ratios and risk ratios, with the result exponentiated back for reporting.
For sparse binary data, where some study arms have very few or zero events, the inverse-variance approach becomes unstable, and the Mantel-Haenszel method is preferred because it handles small cell counts more reliably. The Peto odds ratio is a further option for rare events, valid when events are uncommon and the groups are of similar size. Choosing the right pooling method for the data is part of the method, not an afterthought, and it belongs in the protocol with everything else.
Fixed-effect or random-effects
You then choose a model. A fixed-effect model assumes every study estimates one shared true effect, while a random-effects model assumes the true effect varies and estimates an average of that distribution. Most reviews of real-world evidence default to random-effects because some genuine variation between studies is almost always present. The reasoning behind the choice is worked through in fixed-effect versus random-effects models.
Reading the output
The result is displayed as a forest plot, with each study as a point estimate and confidence interval and the pooled result as a diamond at the bottom. Learning to read one is a skill in itself, covered in forest plot interpretation. Alongside it you report heterogeneity, usually with the I-squared statistic and a between-study variance, explained in heterogeneity in meta-analysis.
Probing the robustness of the result
A single pooled number is never the end. You test whether it holds up by removing influential studies one at a time in a sensitivity analysis, by exploring whether the effect differs across pre-specified groups in a subgroup analysis, and by checking for small-study effects with a funnel plot. Only after these checks, and a GRADE rating of the overall certainty of evidence, is the synthesis ready to write up.
The meta-analysis workflow in order
Pulling the stages together, a defensible quantitative synthesis follows the same sequence every time:
- Confirm the studies are similar enough to pool; if not, write a structured narrative synthesis instead.
- Choose one common effect measure and extract it, with its variance, for every study, converting partially reported studies with the effect size converter.
- Select the model and the tau-squared estimator in advance, then weight studies by inverse variance.
- Compute the pooled estimate and its confidence interval, and quantify heterogeneity with I-squared, tau-squared, and a prediction interval.
- Probe robustness with sensitivity analysis, planned subgroups, and a funnel-plot check for small-study effects.
- Rate the overall certainty with GRADE and report the whole synthesis against PRISMA 2020.
Common mistakes that undermine a pooled estimate
Most flawed meta-analyses fail on a handful of avoidable points. Pooling studies that are too different produces a precise looking number that describes no real population. Mixing incompatible effect measures, such as combining odds ratios with risk ratios, corrupts the pool silently. Switching the model after seeing the heterogeneity turns a planned analysis into a result-driven one. Double-counting by entering the same cohort from two publications inflates precision. And reporting the average while ignoring the prediction interval hides how widely the true effect may vary. Each is prevented by a pre-specified plan and a careful read of the contributing data.
What you need before you start
A meta-analysis cannot rescue a weak review. It needs a complete, reproducible search, duplicate data extraction, and a risk of bias assessment for every study, because the credibility of the pooled estimate depends entirely on the quality of the inputs. The statistical choices, the model, the estimator, and the software, are worked through further in our comparison of meta-analysis tools and our deeper guide to selecting an effect size. Get the inputs right first, and the statistics become the straightforward part.