Synthesis without meta-analysis is the structured combining of study findings when pooling into a single effect size is not appropriate. It is not the absence of a method; it is a set of defined techniques, governed by the SWiM (Synthesis Without Meta-analysis) reporting guideline, that let a review draw defensible conclusions from studies too diverse or too incompletely reported to pool. The difference between a rigorous SWiM and a weak review is precisely whether the synthesis was planned and transparent or whether the authors simply described studies one at a time and hoped a conclusion would emerge.
When pooling is the wrong choice
A meta-analysis is only honest when the studies are similar enough that a single average means something. Three situations rule it out. Clinical diversity arises when populations, interventions, or outcomes differ so much that a pooled number would blend apples and oranges, for example combining a drug trial in children with a surgical trial in adults. Methodological diversity arises when study designs are so mixed that a common effect metric is not meaningful. And incompletely reported effects are the most common practical trigger: studies report a p-value or a direction but no usable effect estimate and variance, so there is nothing to enter into a pooled model. In any of these cases, forcing a meta-analysis produces a precise-looking figure that misleads, and the right move is a structured synthesis without it. Deciding this is a judgement about heterogeneity and comparability, not a failure of effort.
Grouping studies before you synthesise
The first analytic decision in any synthesis without meta-analysis is how to group the studies, and it is as consequential as the choice of method. Grouping by population, intervention type, or outcome domain turns a shapeless pile of studies into comparable clusters within which a direction or a pattern can be read, and it is the non-pooled analogue of a structured subgroup comparison in a meta-analysis. The grouping must be pre-specified and justified, because a reviewer who regroups studies after seeing the results can manufacture an apparent signal, the same danger that makes post-hoc subgrouping suspect in pooled reviews. State the grouping logic, keep it stable, and if you must revise it, say so and explain why. A transparent grouping is what lets a reader follow the argument from individual studies to the overall conclusion without having to take your synthesis on trust.
Structured alternatives that do real work
The core of a good SWiM is choosing a defined synthesis method rather than narrating. Several exist, each with a proper form and known limits. An effect direction plot summarises, for each outcome or study, whether the intervention helped, harmed, or did nothing, using upward and downward arrows, so a reader sees the pattern of direction across a body of evidence that could not be pooled. Vote counting is acceptable only in one form: counting studies by the direction of effect and testing against a null of half the studies favouring each side with a sign test. Counting by statistical significance is not acceptable, because it confuses absence of evidence with evidence of absence and ignores study size; this is the single most common way vote counting is done wrongly. Combining p-values, for instance by Fisher’s method, can test whether there is any effect anywhere in the evidence when effect sizes are unavailable, though it says nothing about magnitude. And an albatross plot back-projects study p-values against sample size to let you read approximate effect magnitudes off contours even when the primary papers never reported them, a clever bridge when data are thin.
Harvest plots for complex evidence
When a review spans many outcomes and mixed designs, a harvest plot extends the effect direction idea into a matrix. Studies are arranged as bars grouped by outcome and by a moderator of interest, with bar height or shading encoding features such as study size or risk of bias, so the reader sees at a glance where the evidence is thick, where it is thin, and where it points. Harvest plots are especially useful in reviews of public health and complex interventions, where the questions are broad and a single pooled estimate was never realistic. They are a presentation method, not a statistical test, so they complement rather than replace the direction-based synthesis, and their shading can encode the study appraisal judgements that a forest plot would otherwise carry.
A worked example of a structured synthesis
Imagine a review of a complex behavioural intervention across fourteen studies that use different outcome measures and mostly report only direction and significance, so a pooled pooled effect estimate is off the table. A rigorous synthesis without meta-analysis groups the studies by outcome domain, tabulates the direction of effect within each, and builds an effect direction plot. Suppose eleven of fourteen studies point toward benefit on the primary domain. A sign test against a null of seven-seven gives a p-value of about 0.06, which you report as suggestive but not conclusive evidence of a consistent effect. You annotate the plot with study size and risk of bias so the reader sees that the three null studies were the largest and least biased, a caveat that a bare count would bury. That combination, a defined method, a formal test of direction, and an honest note about which studies carry the most weight, is a genuine synthesis. It reaches a defensible conclusion without ever pretending the studies could be averaged, and it stands in sharp contrast to a paragraph-per-study description that leaves the reader to do the synthesis themselves.
The nine SWiM items
The SWiM guideline, published by Campbell and colleagues in 2020, is a nine-item checklist that makes a non-pooled synthesis reportable and reproducible. In outline, the items ask you to state: how studies were grouped for synthesis; the standardised metric or direction used; the synthesis method itself; any presentation such as tables or plots; how heterogeneity was investigated; how certainty and robustness were assessed; the criteria used to prioritise results in the write-up; a clear summary of the synthesis findings; and any limitations of the approach. The thread running through all nine is pre-specification: decide the grouping and the method in the protocol and PRISMA reporting before you see the results, so the synthesis cannot be reverse-engineered to fit a preferred conclusion. SWiM sits alongside, not inside, PRISMA 2020, which points to it for reviews that do not pool.
Why “we describe studies one by one” fails peer review
The trap many authors fall into is treating a non-pooled review as licence to abandon method: they write “because the studies were too heterogeneous to pool, we describe each study in turn” and then produce a paragraph per study with no synthesis at all. Reviewers reject this, and rightly, because listing is not synthesising. A reader cannot tell whether the evidence agrees or conflicts, which outcomes are supported, or how strong the overall signal is. The failure is not that the authors chose not to pool; it is that they chose no method to replace pooling. A defensible review names its synthesis approach (effect direction, vote counting by direction, an albatross or harvest plot), applies it consistently, reports it against the nine SWiM items, and still reaches a structured conclusion with a stated certainty. This is the same discipline as a formal narrative synthesis compared with meta-analysis, and it draws on the transparent data coding used in qualitative reviews to keep grouping decisions auditable.
The limits of each alternative method
Every structured method buys transparency at a price, and a good review states the price. Vote counting by direction ignores the size of each effect and the precision of each study, so it can call a body of evidence “consistent” when the consistent direction hides trivially small effects; it answers only whether there is any effect, not how big. Combining p-values by Fisher’s or Stouffer’s method shares that blindness to magnitude and is sensitive to a single very small p-value, so one large study can dominate the combined result. Effect direction plots and harvest plots are presentational: they organise the evidence beautifully but perform no inference, so a reader could still draw an unwarranted conclusion from a striking-looking display. Albatross plots recover only approximate magnitudes and depend on the assumed analysis behind each p-value. None of this argues against the methods; it argues for naming their limitation next to their result, exactly as a pooled review reports its statistical heterogeneity and its risk of bias in the evidence base. A synthesis that pretends its method has no blind spot is as fragile as one that forced an inappropriate pooled average.
Grading certainty when you did not pool
A common worry is that a non-pooled review cannot support a certainty rating, but this is false. GRADE can be applied to a synthesis without meta-analysis, with some adaptation: the domains of risk of bias, inconsistency, indirectness, imprecision, and publication bias still apply, though imprecision and inconsistency are judged from the pattern of directions and the spread of results rather than from a confidence interval around a pooled effect. Where between-study variation would normally be quantified, you reason instead about how consistently the studies point the same way and how much they vary in magnitude where magnitude is known. The result is a certainty judgement, often lower than a pooled review would reach because the evidence was too diverse to combine, but a judgement nonetheless. Feeding that into a summary of findings table keeps the review comparable with pooled ones and prevents a reader from mistaking “we could not pool” for “we could not conclude”. A formal robustness check on the grouping decisions, repeating the synthesis with an alternative grouping of outcomes, strengthens the certainty argument further.
Getting the synthesis right from the protocol
The strongest SWiM is planned before the data arrive. At the protocol stage you should state the circumstances under which you will not pool, the grouping of studies you expect, and the synthesis method you will use if a meta-analysis proves inappropriate, so the decision is principled rather than convenient. During the review you apply that plan, present the evidence with the chosen structured method, investigate heterogeneity, and grade the certainty of the evidence exactly as a pooled review would. The recurring mistakes are three: defaulting to vote counting by significance when only direction is defensible; describing studies one by one with no synthesis method at all; and failing to report against the SWiM items, which leaves the synthesis looking improvised. Plan the method, apply it consistently, and report it fully, and a synthesis without meta-analysis carries every bit as much weight as a pooled one.