Subgroup analysis and meta-regression are the two main ways a meta-analysis explores why studies disagree. A subgroup analysis splits studies into categories, such as adults versus children, and compares the pooled effect in each group. Meta-regression goes further by modelling the pooled effect size as a function of one or more study-level characteristics, including continuous ones like average dose or year of publication. Both are tools for investigating heterogeneity, not for proving cause.

Why explaining heterogeneity matters

When a forest plot shows results scattered far more than chance allows, a single pooled number can mislead. High heterogeneity signals that the true effect may differ across settings, populations, or designs. Rather than report one average that fits nobody, you ask what explains the spread. That investigation is what subgroup analysis and meta-regression are for, and the candidate explanations should be named in the protocol before the data are seen.

How subgroup analysis works

A subgroup analysis partitions the included studies by a categorical variable, pools each group separately, then tests whether the subgroup effects differ. The key output is a test for subgroup differences, which asks whether the gap between groups is larger than sampling error would produce. A common error is reading a significant effect in one subgroup and a non-significant one in another as proof they differ; only the formal interaction test answers that.

A worked subgroup example

Suppose a meta-analysis of an antidepressant pools twelve trials and finds a modest benefit overall, with high heterogeneity. You pre-specified baseline severity as a moderator, so you split the trials into those recruiting moderate depression and those recruiting severe depression. The severe subgroup shows a standardised mean difference of 0.55, the moderate subgroup 0.12. It is tempting to announce that the drug works in severe but not moderate disease. The test for subgroup differences asks the proper question: is the gap between 0.55 and 0.12 larger than sampling error alone would produce? If that interaction test returns a non-significant result, the apparent difference is unreliable, no matter how convincing the two separate forest plots look. Reporting the interaction test, not the two within-group significance verdicts, is the single most important discipline in subgroup analysis.

Pre-specified versus post-hoc subgroups

Subgroups defined in advance, based on a plausible mechanism, carry far more weight than ones invented after seeing the results. Post-hoc subgroups invite data dredging, where you keep slicing until something looks significant. The arithmetic is unforgiving: testing ten independent subgroups at the conventional five percent threshold gives roughly a 40 percent chance of at least one false positive purely by chance. Keeping the list short and pre-specified, reporting every subgroup you examined, and being cautious about claims from any unplanned split protects the review from this trap, much as a registered PROSPERO record protects the rest of the method. The same caution that drives a strict set of eligibility rules applies to the analysis plan.

How meta-regression works

Meta-regression treats each study as a data point and fits a line, or surface, relating the effect size to a moderator variable. A continuous moderator, such as mean participant age or treatment duration, is something subgroup analysis cannot handle without crude cut-points. The model returns a slope: how much the effect changes per unit of the moderator, with a confidence interval around it. As with any regression, you need enough studies, and the usual guidance is to avoid fitting more than one moderator per handful of studies.

Reading the residual heterogeneity

A good moderator reduces the leftover, or residual, heterogeneity. Software reports an R-squared analog, the percentage of between-study variance the moderator explains, and the residual tau-squared that remains after fitting it. If a moderator absorbs most of the spread, leaving a small residual, it is a strong candidate explanation; if the residual tau-squared stays high, the real driver is still unmeasured. You can quantify the starting heterogeneity with the heterogeneity calculator before deciding whether modelling it is even worthwhile, and read the slope alongside the random-effects assumptions that almost always apply to a meta-regression, since residual heterogeneity is rarely zero.

How many studies you really need

Meta-regression is a regression, and regressions overfit when there are too few data points per predictor. The widely cited rule of thumb is at least about ten studies per moderator, so a single continuous moderator wants roughly ten studies, two moderators want twenty, and so on. With fewer, a moderator can appear to explain the spread purely by chance because the line has enough freedom to pass near any scatter of a handful of points. This is why a review of six trials should almost never report a meta-regression as a finding; at most it is a hypothesis to test in future work. The same constraint makes a categorical subgroup analysis the more honest choice when studies are scarce, because it asks a simpler question of the data.

The ecological fallacy and other cautions

Both methods relate study-level averages to outcomes, so a relationship seen across studies need not hold within individuals, a trap known as the ecological fallacy. They are also observational: studies are not randomised to subgroups, so a difference may reflect confounding rather than the moderator itself. Treat both as hypothesis-generating, and weigh the findings alongside the risk of bias in the contributing studies and the overall certainty of evidence.

Common mistakes to avoid

Five errors recur in submitted reviews, and each is avoidable:

  • Comparing significance instead of testing interaction: declaring a difference because one subgroup is significant and another is not, without the formal test for subgroup differences.
  • Too many moderators for too few studies, which overfits and produces explanations that will not replicate.
  • Post-hoc subgroups presented as planned, which inflates the false-positive rate and misleads the reader about how the hypothesis arose.
  • Reading a study-level slope as an individual effect, the ecological fallacy, when the relationship may not hold within patients at all.
  • Treating a subgroup difference as causal, when studies were never randomised to subgroups and confounding by other study-level features is unaccounted for.

Where these fit in the analysis plan

Subgroup analysis and meta-regression come after the main pooled estimate, once you have run the meta-analysis and seen genuine heterogeneity worth explaining. They sit alongside sensitivity analysis, which tests robustness rather than explanation, and they feed the inconsistency judgement when you rate the overall certainty of evidence. Used together and pre-specified, they turn an unexplained spread of results into a structured account of when and for whom an effect is larger or smaller. If you want these analyses designed and reported so they hold up to peer review, our meta-analysis service plans and runs the synthesis end to end.