Egger’s test is a linear regression that quantifies funnel plot asymmetry in a meta-analysis. It regresses each study’s standardised effect, the effect divided by its standard error, on its precision, the reciprocal of that standard error. If the fitted line passes close to the origin the funnel is symmetrical; if the intercept departs significantly from zero, small and large studies are giving systematically different answers, the statistical fingerprint of a small-study effect.

The regression that sits behind the plot

A funnel plot is a scatter of effect against precision that a reader eyeballs for a missing corner. Egger turned that visual impression into a number. In the original 1997 formulation you regress the standardised effect (effect over standard error) on precision (one over standard error), weighting by inverse variance. An equivalent and more common phrasing regresses the raw effect on its standard error, so the slope carries the clinical effect and the intercept carries the asymmetry. Either way, the intercept is the quantity of interest: under symmetry its expected value is zero, because small imprecise studies should scatter evenly above and below the pooled effect. A non-zero intercept means the small studies are pulled to one side, which is exactly the pattern a gap of unpublished null results would create.

Reading the intercept and the p-value

The output you report is the estimated intercept, its confidence interval, and a p-value for the null hypothesis that the intercept is zero. Because these tests are underpowered, many methodologists, including the Cochrane Handbook, suggest judging significance at p < 0.10 rather than the usual 0.05. A significant intercept says the funnel is asymmetrical to a degree unlikely to be chance; it does not say by how much your pooled estimate is wrong, and it does not name the cause. The sign of the intercept indicates the direction of the asymmetry, so it tells you whether small studies are exaggerating or understating the effect relative to the large ones. Report the number, not just the verdict, so a reader can see how borderline the finding is.

The ten-study minimum, and why it matters

The single most abused feature of Egger’s test is that people run it on a handful of studies. With few studies the test has almost no power, so a non-significant result is uninformative rather than reassuring, and a significant one is unstable. The widely cited rule of thumb is a minimum of ten studies before a formal asymmetry test is worth interpreting, and even then the result is a prompt, not a proof. Below that threshold the honest sentence in your methods is that the number of studies was too small to assess funnel plot asymmetry reliably, which is a defensible position and far stronger than quoting a meaningless p-value. This same numerical humility runs through good robustness checking of a pooled estimate.

Why the test flags small-study effects, not bias itself

A significant intercept is evidence of a small-study effect, the tendency for smaller trials to report larger effects. Suppressed null studies are one explanation, but not the only one. Genuine heterogeneity, where smaller trials recruit sicker patients or deliver a more intensive intervention, produces the same asymmetry without any missing studies. So can poorer methods in small trials, or the mathematical artefact discussed below. This is why Egger’s test is a detector of asymmetry and not a detector of publication bias, which owns the broader topic of missing evidence. Treat a significant result as a reason to investigate whether a subgroup or meta-regression explanation fits before you conclude that studies were withheld.

The problem with binary outcomes

When the effect is an odds ratio or risk ratio, the effect estimate and its standard error are mathematically linked, because both depend on the same event counts. That structural correlation makes the original Egger’s test reject too often, inflating the false-positive rate even when nothing is missing. Two modifications were built to fix this. The Harbord test uses a score statistic and its variance, which breaks the artificial link and controls the error rate for log odds ratios. The Peters test regresses the effect on the inverse of the total sample size, again sidestepping the count dependence, and behaves well for proportions and rare events. For continuous outcomes measured as a mean difference the original Egger regression is usually fine; the modifications matter most when your effect metric is built from the same numbers as its own variance.

A worked example of the regression

Imagine twelve trials of an intervention, with effect sizes ranging from a small standardised mean difference in the large trials to a much larger one in the smallest. You compute each trial’s standardised effect and its precision, then fit the weighted regression. Suppose the line has an intercept of 1.8 with a 95 percent confidence interval of 0.4 to 3.2 and a p-value of 0.02. Because the interval excludes zero and the p-value is below the 0.10 threshold, you report clear evidence of funnel plot asymmetry: the smaller, less precise trials are systematically reporting larger effects than the big ones. The positive sign tells you the small studies are inflating the pooled effect, so the true effect is probably nearer the estimate from the large trials alone. Notice what the number does not tell you: it does not say whether four null trials were suppressed or whether the small trials simply used a more responsive population. That interpretive step is yours, and it should be argued from the clinical context, not read mechanically off the p-value.

Where Egger’s test sits among the asymmetry tests

Egger’s test is one member of a family. Its oldest companion is Begg’s rank-correlation test, which correlates the standardised effects with their variances using Kendall’s tau; it makes fewer parametric assumptions but has notably lower power, so it rarely detects asymmetry that Egger misses and often fails to confirm asymmetry that Egger flags. The Thompson and Sharp test extends Egger by allowing for between-study heterogeneity in the regression, which is sensible when the studies are not estimating one common effect, and it feeds directly into the way you should think about a random-effects model. For the binary-outcome situation, Harbord and Peters (covered above) are the recommended defaults. The practical guidance is to pre-specify one primary test suited to your effect metric and study count, report it in full, and treat any secondary tests as supportive rather than as a menu you sample until one turns significant. Running five tests and quoting the one that agrees with your prior is the multiple-testing trap that undermines the whole exercise.

Running it in R and Stata

In R with the metafor package you fit your model with rma() and then call regtest(), which performs Egger’s regression and lets you switch the predictor to sample size for the Peters variant. The meta package offers metabias() with a method.bias argument that selects Egger, Harbord, Peters, or Begg. In Stata the modern meta bias command takes egger, harbord, or peters as options after you declare your data with meta set or meta esize. Whichever route you take, pair the test with a contour-enhanced funnel plot so the reader sees the shape and the statistic together. You can generate both quickly in our funnel plot and asymmetry calculator before committing to the full model, and cross-check the pooled figure in the meta-analysis calculator.

The effect metric can manufacture asymmetry

A subtlety that catches out even experienced reviewers is that the shape of the funnel, and therefore the result of the test, can depend on which effect metric you chose to pool. The same body of trials can look symmetrical when analysed as a risk difference and asymmetrical as an odds ratio, because the mathematical relationship between the effect and its standard error differs across metrics. This is a further reason a significant intercept is not proof of missing studies: it may be an artefact of the scale. A sensible guard is to check whether the asymmetry persists under a plausible alternative metric before you build a conclusion on it, and to present the primary forest plot on the scale you pre-specified rather than the one that happens to look cleanest. When asymmetry appears on one scale and vanishes on another, the honest report says so and treats the finding as inconclusive rather than picking the version that supports the narrative.

Reporting Egger’s test properly

A defensible write-up states four things: how many studies fed the test, which test you used and why (naming Harbord or Peters for binary data), the estimated intercept with its confidence interval and p-value, and your interpretation in terms of small-study effects rather than a bald claim of publication bias. Pre-specify all of this in the PRISMA 2020 reporting so the analysis cannot look like a fishing expedition. If the test is significant, feed that judgement into the publication-bias domain of the GRADE certainty rating and consider a sensitivity analysis such as trim and fill to gauge how far the conclusion could move. What you must not do is present the intercept as the last word: it is one diagnostic among several, most persuasive when it agrees with the funnel plot and a wide, well-documented search.

Why a symmetrical funnel is not a clean bill of health

It is tempting to read a non-significant Egger’s test as confirmation that no publication bias exists, but the logic does not run that way. The test has limited power, so with a modest number of studies it can easily fail to detect real asymmetry, and a symmetrical funnel is consistent with several forms of bias it cannot see at all. Outcome reporting bias, where a published study quietly drops its disappointing endpoints, leaves the funnel of the reported outcome perfectly symmetrical while still distorting the evidence. So does the wholesale suppression of a set of studies that happen to be spread evenly across effect sizes. The honest interpretation of a non-significant result is therefore modest: it found no evidence of funnel plot asymmetry, given the studies available, which is not the same as evidence that the body of evidence is complete. That distinction, absence of evidence versus evidence of absence, is exactly the one a certainty assessment is built to respect, and it is why a wide search remains the primary defence no matter what the test returns.

Common mistakes with Egger’s test

The recurring errors are easy to list and easy to avoid. Reviewers run the test on fewer than ten studies and report the p-value as if it meant something. They use the original Egger regression on odds ratios and mistake its inflated false-positive rate for real asymmetry, when Harbord or Peters was the correct choice. They read a significant intercept as proof of suppressed studies without ruling out heterogeneity or study quality. And they quote the verdict without the number, hiding how borderline the result was. Naming the test, stating the study count, and interpreting the intercept as a small-study signal removes all four in a single honest paragraph.