Publication bias is the distortion that arises when studies with statistically significant or favourable results are more likely to be published, and therefore easier to find, than studies with null or unfavourable results. In a meta-analysis, this means the studies a reviewer can locate may overstate the true effect size, because part of the evidence stays in file drawers and conference abstracts that never appear.
Why the missing studies matter most
A pooled estimate is only as honest as the set of studies feeding it. If small studies that found nothing were quietly never written up, averaging the survivors inflates the result. This is why publication bias is treated as a threat to the whole synthesis rather than a flaw in any one trial, and why it is one of the five reasons GRADE can downgrade the certainty of an outcome. It is a different problem from the per-study risk of bias assessment: the individual studies may each be sound, yet the body of evidence is still skewed by what is absent.
How reviewers detect it
The funnel plot
The first line of investigation is a funnel plot, which charts each study’s effect against its precision. With no bias, the points form a roughly symmetrical inverted funnel; a gap in one bottom corner, where small unfavourable studies should sit, suggests they are missing. Reading that shape correctly takes practice, which we cover in detail in funnel plot interpretation.
Formal tests for asymmetry
Visual judgement is subjective, so reviewers add a statistical test of funnel plot asymmetry. Egger’s test regresses each study’s standardised effect (the effect divided by its standard error) on its precision (the reciprocal of the standard error); an intercept that differs significantly from zero is the signal of asymmetry. Begg’s test takes a rank-correlation approach instead, correlating the standardised effects with their variances, and tends to have lower power than Egger’s. Both need enough studies to be meaningful, typically around ten or more, and a significant result points to small-study effects rather than proving publication bias outright. For binary outcomes where the effect and its variance are mathematically linked, the Harbord or Peters modifications reduce the false-positive rate that the original Egger’s test can show.
Trim-and-fill and the fail-safe N
Where a test flags asymmetry, the trim-and-fill method estimates how many studies appear to be missing, imputes mirror-image points to restore symmetry, and re-pools to give an adjusted estimate. The gap between the original and adjusted effect is a rough sensitivity check: if the adjusted effect still clears the line of no effect, the conclusion is more robust to suspected missing studies. Treat the imputed studies as a thought experiment, not as real data, because trim-and-fill performs poorly under genuine heterogeneity. An older device, Rosenthal’s fail-safe N, counts how many null studies it would take to overturn a significant pooled result; it is largely deprecated because it ignores effect sizes and can give reassuringly large numbers even when bias is plausible. The contemporary preference is a contour-enhanced funnel plot read alongside a regression test, which you can generate in our funnel plot and asymmetry calculator.
A worked example of how missing studies inflate an effect
Suppose eight published trials of an intervention pool to a risk ratio of 0.70, a thirty percent reduction in events, with a tight confidence interval that excludes 1. A trim-and-fill analysis suggests four small null trials are missing from the lower-right of the funnel. Filling them in pulls the pooled risk ratio back toward 0.85 and widens the interval so it now brushes 1. The headline did not vanish, but a confident “thirty percent reduction” has become a modest, uncertain effect. That movement, not the original tidy figure, is what an honest discussion reports, and it is exactly the kind of fragility a robustness check on the pooled estimate is designed to expose.
Why asymmetry is not always publication bias
A lopsided funnel has several possible causes, and conflating them is a common mistake. True heterogeneity, where smaller studies genuinely differ in their populations or intervention intensity, can produce the same shape. So can poorer methods in smaller trials, a form of small-study effect that overlaps with bias. A careful review treats asymmetry as a prompt to investigate, examining whether a subgroup or meta-regression explanation fits before concluding that studies are missing.
The forms publication bias actually takes
The phrase is a shorthand for several distinct mechanisms, and naming the right one sharpens both detection and defence. Classic publication bias is the wholesale non-publication of null studies. Outcome reporting bias is subtler: a study is published, but the disappointing outcomes are dropped or relegated while the favourable ones are promoted, which is why comparing a paper to its registered protocol matters. Time-lag bias sees positive results published faster, so a review run too early oversamples them. Language bias and database bias arise when significant findings are more likely to appear in English-language or widely indexed journals, so restricting the search to convenient sources quietly tilts the evidence. Each of these is a separate hole that a broad comprehensive search strategy is built to plug.
Guarding against it from the start
The strongest defence is built before the analysis. A thorough grey literature search for theses, preprints, and trial registries recovers studies that never reached a journal, and a registered protocol on PROSPERO commits you to reporting all pre-specified outcomes regardless of how they turn out. Detection methods catch what slips through, but a wide search and an honest protocol stop much of the bias entering in the first place. Practical steps that materially reduce the risk include:
- Search trial registries such as ClinicalTrials.gov and the World Health Organisation portal for completed-but-unpublished studies, then attempt to retrieve their results.
- Use citation chasing in both directions to surface studies the database search missed, as covered in backward and forward citation searching.
- Avoid language restrictions where feasible, and translate or screen non-English records rather than excluding them wholesale.
- Contact authors directly for unreported outcomes or analyses you can see were collected.
Common mistakes when handling publication bias
Three errors recur in submitted reviews. The first is running a formal test on too few studies, then reporting its p-value as if it were trustworthy; with five or six studies the test is essentially uninformative and the honest statement is that asymmetry could not be assessed. The second is treating asymmetry as proof of suppression without ruling out the innocent explanations covered above, which overstates a problem that may be ordinary heterogeneity. The third is reporting trim-and-fill as a correctionrather than a sensitivity check, presenting the adjusted estimate as the true effect when it is only a guess about what missing studies might have shown. A defensible review pre-specifies its bias assessment in the protocol, states how many studies it had, names the test, and feeds the result into the publication-bias domain of the GRADE judgement rather than burying it.