Effect sizes in meta-analysis are the standardised numbers that put every study on a common scale so their results can be pooled. An effect size expresses the direction and magnitude of a relationship, such as a treatment’s benefit, in a way that does not depend on the original measurement units or sample size. Common metrics include the odds ratio, risk ratio, mean difference, and the standardised mean difference, and the right choice depends on whether the outcome is binary or continuous.

Why the choice of metric shapes the whole synthesis

Picking the wrong metric quietly distorts every result that follows. A meta-analysis pools effect estimates and their standard errors, so if studies report the same finding on incompatible scales, the pooled number is meaningless. The metric also drives how you read heterogeneity and how the result is communicated to clinicians, who think in absolute terms more readily than ratios. Settling on the metric belongs in the protocol, long before any data are extracted.

Effect sizes for binary outcomes

When the outcome is a yes-or-no event, such as death, relapse, or recovery, the candidates are the odds ratio, the risk ratio, and the risk difference. The odds ratio is mathematically convenient and the default in many models, but it is easily misread as a risk ratio when the event is common. Our guide on odds ratio versus risk ratio works through when each is appropriate. Because these ratios are skewed, they are pooled on the log scale and converted back for reporting.

Turning raw counts into a ratio

From a two-by-two table of events and non-events in each arm you can compute any of the binary metrics directly. Suppose a trial records 20 events in 100 treated patients and 40 events in 100 controls. The risk ratio is 0.20 divided by 0.40, or 0.50, a halving of risk. The odds ratio compares 20/80 against 40/60, giving 0.375, a more extreme-looking number for the same data, which is precisely why the two are easy to confuse when the event is common. The odds ratio and risk ratio calculator does this and returns a confidence interval, which is the input a pooling routine actually needs. The translation between a relative and an absolute reduction matters to readers, and the number needed to treat is often the clearest way to communicate it.

Why ratios are pooled on the log scale

A ratio of one means no effect, but the scale is asymmetric: a doubling of risk is a ratio of 2, while a halving is 0.5, and those are not equidistant from 1 on a linear scale even though they are mirror-image effects. Taking the natural logarithm fixes this, because the log of 2 and the log of 0.5 are equal and opposite. Meta-analyses therefore pool the log odds ratio or log risk ratio, whose sampling distribution is far closer to symmetric and whose standard error has a simple closed form, then exponentiate the pooled result and its interval back to the natural ratio for reporting. Forgetting this step is a classic error that distorts both the pooled estimate and its confidence interval.

Effect sizes for continuous outcomes

When studies measure something on a scale, such as a depression score or blood pressure, the choice is between the mean difference and the standardised mean difference. Use the raw mean difference when every study used the same instrument, because it keeps the natural units. Use the standardised mean difference when studies measured the same construct with different scales, since dividing by the pooled standard deviation removes the units entirely.

Standardising a mean difference

The standardised mean difference expresses the gap between two groups in standard-deviation units. The raw form, dividing the difference in means by the pooled standard deviation, is Cohen’s d, but it is biased upward in small samples, so most meta-analyses apply the Hedges’ g correction, which multiplies by a factor a little below one that shrinks toward d as the sample grows. A worked case: two groups with means of 52 and 48 and a pooled standard deviation of 10 give a Cohen’s d of 0.40, a moderate effect, regardless of whether the scale ran 0 to 100 or 0 to 20, which is exactly what makes it poolable across instruments. You can compute it from group means, standard deviations, and sample sizes with the standardised mean difference calculator. A value near 0.2 is conventionally small, 0.5 moderate, and 0.8 large, though those labels are rough guides rather than rules, and a small standardised effect can still be clinically important when the outcome is serious.

The variance is as important as the point estimate

Every effect size travels with a standard error, and the pooling routine needs both. The standard error is what becomes the inverse-variance weight, so an effect entered without a usable variance cannot contribute to the pool at all. This is why extraction must capture not just the mean difference or ratio but the figures that yield its variance: standard deviations and group sizes for continuous outcomes, event counts and totals for binary ones. When a paper reports only a confidence interval or a p-value, the standard error can usually be reconstructed, and the confidence interval calculator helps recover it. A study whose variance cannot be recovered, even after contacting the authors, is the one most likely to be lost from the synthesis.

Converting between effect sizes

Real reviews rarely find every study reporting the same statistic. One trial gives a correlation, another an odds ratio, a third only a t-value and sample sizes. To pool them you convert everything into one target metric using established formulas. The effect size converter handles the common transformations, for example between a correlation, an odds ratio, and a standardised mean difference, so partially reported studies are not wasted. When a study omits the numbers you need, the next step is to record what the study did report and reconstruct the rest where the formulas allow.

From effect sizes to a pooled estimate

Once every study carries the same metric and a standard error, the results are combined into a single weighted estimate. Whether you use a fixed-effect or random-effects model changes how studies are weighted, and the pooled effect size with its confidence interval is what you display on a forest plot. You can try the full pooling step in the meta-analysis calculator, which reports the summary effect alongside an I-squared heterogeneity figure.

Common mistakes with effect sizes

The frequent errors are worth listing plainly:

  • Mixing incompatible metrics in one pool, such as combining odds ratios with risk ratios, which corrupts the summary silently.
  • Forgetting the log scale for ratio measures, which distorts both the pooled estimate and its interval.
  • Ignoring the sign when outcome scales run in opposite directions, so a higher score means improvement in one study and deterioration in another.
  • Treating a standardised mean difference as a clinical unit, when it is expressed in standard deviations and must be back-translated to a familiar scale for interpretation.
  • Entering an effect without its variance, which leaves the study unweighted and effectively excluded from the pool.

Each of these inflates or hides real differences, so a careful piloted extraction form that records the metric and its variance for every study is the best defence. When effect sizes arrive in a dozen incompatible formats and need harmonising before they can be pooled, our meta-analysis service extracts, converts, and combines them to a reproducible standard.