A network meta-analysis compares three or more interventions at once by combining direct evidence, from trials that compared two treatments head to head, with indirect evidence, inferred through a common comparator. If trials compare treatment A with C and others compare B with C, the network can estimate the relative effect of A versus B even when no trial compared them directly. The output is a coherent set of relative effects across every pair of treatments and a ranking of how each is likely to perform, all from a single connected analysis.

How a network extends the standard pairwise approach

A conventional, pairwise meta-analysis combining study results answers one question: how does treatment A compare with treatment B. Clinical decisions rarely come down to two options, though. A clinician choosing among five drugs for the same condition needs them all on one scale, and the trials almost never line up neatly: some compared drug 1 with placebo, others compared drug 3 with drug 1, and so on. Network meta-analysis, also called multiple treatments meta-analysis or mixed treatment comparison, links these fragments into a single evidence network in which every treatment is a node and every direct comparison is an edge. Because it borrows strength across the whole network, it can produce estimates and a complete ranking that no single trial or pairwise analysis could deliver.

Direct, indirect, and mixed estimates

The core mechanic is the indirect comparison. Suppose the direct evidence puts A versus C at a log odds ratio of -0.40 and B versus C at -0.10. The indirect estimate of A versus B is the difference, -0.30, anchored through the shared comparator C. Where both direct and indirect evidence exist for the same pair, the network combines them into a more precise mixed estimate, weighting each by its precision in the same spirit as the inverse-variance weighting used in pooled analyses. This borrowing of strength is the source of the method’s power and also of its fragility, because an indirect estimate is only as sound as the assumptions that licence it.

The two assumptions that make or break a network

Everything rests on two linked assumptions. Transitivity is the clinical and methodological assumption that the studies forming each comparison are similar enough in their participants, settings, doses, and definitions that the common comparator behaves the same way across them. If the placebo arms in the A versus C trials enrolled mild cases while those in the B versus C trials enrolled severe cases, the indirect A versus B estimate is contaminated by that difference. Transitivity cannot be tested directly; it is judged by tabulating the distribution of effect modifiers across comparisons before any pooling, and differences here are a network-level form of the heterogeneity that troubles any synthesis.

Consistency (sometimes called coherence) is the statistical counterpart: the requirement that the direct and indirect estimates for the same comparison agree. It can be tested, for example by node-splitting, which separates the direct and indirect evidence for each comparison and checks whether they conflict, or with a design-by-treatment interaction model. Inconsistency is a red flag that transitivity has failed somewhere in the loop, and it must be investigated rather than averaged away. Transitivity is the assumption; consistency is the observable consequence you can interrogate.

Network geometry

Before any estimation, plot the network geometry: a diagram with a node per treatment, sized by how many participants received it, and an edge per direct comparison, weighted by how many trials inform it. The shape tells you a great deal. A well-connected network with many closed loops has plenty of direct and indirect evidence to cross-check. A star-shaped network, where every treatment connects only to a common comparator such as placebo and to nothing else, has no closed loops, so consistency cannot be checked at all and the whole analysis leans entirely on transitivity. Sparse connections and a single trial bridging two halves of the network are warnings that the indirect estimates will be fragile.

Ranking treatments with SUCRA

Because a network estimates every treatment’s relative effect, it can rank them, and the most common summary is the surface under the cumulative ranking curve, abbreviated as SUCRA. SUCRA expresses, as a percentage, how much of the “ideal” top ranking a treatment achieves on average across the uncertainty in the estimates: a value near 100 percent means the treatment is almost always ranked best, and a value near zero means it is almost always ranked worst. The crucial caution is that ranking is not the same as a clinically important difference. A treatment can top the SUCRA ordering while its advantage over the second-ranked option is tiny and its confidence intervals on the relative effect overlap heavily. Always read the rankings alongside the effect estimates and their intervals, never on their own.

When a network is appropriate, and when it is not

A network meta-analysis is appropriate when there are multiple competing interventions for the same condition and outcome, when the trials form a connected network (every treatment must link to the rest, directly or indirectly), and when transitivity is clinically plausible. It is the wrong tool when the evidence base splits into disconnected sub-networks that share no comparator, when the populations or outcome definitions differ so much that transitivity is indefensible, or when a simple pairwise question is all that is being asked. In those cases a standard meta-analysis, or a structured narrative synthesis where pooling is unsafe, is the honest choice. The certainty of each comparison should be rated with the network extension of the GRADE approach to certainty, which incorporates both direct and indirect contributions and any incoherence.

How many studies you need, and the disadvantages

There is no magic minimum. Technically a single trial per comparison can connect a network, but a network resting on single trials and a handful of thin indirect links produces unstable, wide estimates and cannot be checked for consistency. What matters is not a raw count but a connected network with enough closed loops to test coherence and a credible case for transitivity, much as the right question for a pairwise pool is whether the studies are similar enough to combine rather than how many there are. The disadvantages are real and follow from the assumptions: results are vulnerable to intransitivity when populations differ across comparisons; inconsistency between direct and indirect evidence can invalidate estimates; rankings are easy to over-interpret; and the analysis is statistically and conceptually demanding, often run in a Bayesian framework that requires careful prior specification and convergence checking. Exploring why effects differ across the network through meta-regression on effect modifiers is often essential rather than optional. Handled with discipline, a network meta-analysis answers questions no single trial can; handled carelessly, it manufactures false precision across a dozen comparisons at once.