A prevalence meta-analysis pools the proportions reported across studies into a single estimate of how common a condition or characteristic is in a population. Unlike a meta-analysis of a comparison, there is no control group and no effect size contrast; the quantity being combined is a single proportion from each study. That sounds simpler, but raw proportions misbehave statistically near the boundaries of 0 and 1, so the method turns on variance-stabilising transformations and on handling the extreme heterogeneity these reviews almost always show.
Why raw proportions misbehave near 0 and 1
A proportion is bounded between 0 and 1, and its sampling variance depends on the proportion itself, shrinking toward zero as the value approaches either boundary. Two consequences follow. First, a study reporting a prevalence of 2 percent has an artificially tiny variance, so in an inverse-variance pooling it receives huge weight and can dominate the result. Second, the usual assumption that the estimate is normally distributed breaks down: confidence intervals computed on the raw scale can run below 0 or above 1, which is nonsensical. Studies clustered near 0 (a rare disease) or near 1 (a near-universal exposure) are where this bites hardest. The fix is not to pool the raw proportions at all, but to transform them onto a scale where the variance is roughly constant and the normal approximation holds, pool there, and transform back.
Logit and double arcsine transformations
Two transformations dominate. The logit transformation maps a proportion onto the log-odds scale, which is unbounded and familiar from logistic regression; it is well understood, the pooled result back-transforms cleanly to a proportion, and its weakness is that it still struggles when many studies report exactly 0 or 1, which need a continuity correction. The Freeman-Tukey double arcsine transformation stabilises the variance more effectively, especially for proportions close to the boundaries, and for years was the recommended default. Its Achilles heel is the back-transformation: converting the pooled double arcsine value back to a proportion requires an assumed sample size, and Schwarzer and colleagues showed in 2019 that this step can produce seriously misleading pooled proportions, sometimes even outside the range of the observed studies. The current debate has shifted the balance: many methodologists now prefer the logit for its honest and interpretable back-transformation, or move to a generalised linear mixed model that models the counts directly and sidesteps transformation altogether. The safe stance is to pre-specify your transformation, and to report the pooled estimate on a scale the reader can trust.
Random effects as the default
In a prevalence review the fixed-effect model is almost never plausible. A fixed-effect model assumes every study is estimating the same underlying proportion, but prevalence genuinely differs between countries, age bands, diagnostic criteria, and years, so the true value varies from study to study by design. The random-effects model, which treats each study as estimating its own true prevalence drawn from a distribution, is therefore the default, and the pooled figure should be read as an average across settings, not a single universal rate. Even then, the average can be less useful than the spread, which is why a prediction interval, the range within which a new study’s true prevalence is expected to fall, is often the most honest headline number a prevalence review can report.
Extreme heterogeneity is normal, so plan for it
Heterogeneity in prevalence reviews is routinely enormous, with an I-squared above 95 percent being common rather than alarming, because true prevalence really does vary between populations. Chasing a low I-squared is the wrong goal; the right goal is to explain and describe the variation. Three tools do the work. A prediction interval communicates the realistic range instead of pretending the average is precise. Subgroup analysis and meta-regression test whether region, age, diagnostic definition, or study year account for the spread, turning uninterpretable heterogeneity into structured findings. And a frank narrative account of why prevalence differs keeps the review honest where statistics run out. Reporting a single pooled percentage with a 95 percent I-squared and no further comment is the classic error reviewers flag; the variation is the finding, not a nuisance to be buried.
A worked example
Suppose twenty surveys estimate the prevalence of a condition, reporting values from 3 percent to 41 percent. Pooling the raw proportions under inverse-variance weighting would let the largest, most precise survey dominate and could return a nonsensical interval. Instead you transform each proportion to the logit scale, pool under random effects, and back-transform to a pooled prevalence of, say, 18 percent with a 95 percent confidence interval of 14 to 23 percent. The I-squared is 98 percent. Reporting only the 18 percent would be misleading, because the prediction interval runs from 5 to 44 percent, telling you that a new survey in a new setting could plausibly find almost any value in that band. That prediction interval, not the tidy point estimate, is the honest headline, and it immediately motivates a meta-regression on region and diagnostic criteria to find out what drives the spread. A review that stops at “pooled prevalence 18 percent” has hidden its most important result.
Zero cells and near-universal proportions
Two boundary situations need explicit handling. When a study reports zero events, so an observed proportion of exactly 0, the logit is undefined, and the traditional fix is a continuity correction that adds a small constant such as 0.5 to the counts. Continuity corrections are a crude patch that can bias the pooled estimate, especially when many studies have zero or very low counts, which is a strong argument for a generalised linear mixed model that models the binomial counts directly and needs no correction at all. The mirror-image problem arises with near-universal proportions close to 1, where the same instability appears at the top boundary. Whichever route you take, run a sensitivity analysis that repeats the pooling under an alternative transformation or model, and report whether the estimate is stable. If the double arcsine and the logit give materially different pooled proportions, that disagreement is itself a finding to disclose, echoing the wider caution reviewers apply to bias in the body of evidence.
Appraising the studies with JBI
Prevalence studies are not randomised trials, so the usual trial-focused appraisal tools do not fit. The Joanna Briggs Institute (JBI) critical appraisal checklist for studies reporting prevalence data is the standard instrument. Its nine items probe the things that bias a prevalence estimate: whether the sample frame represented the target population, whether participants were sampled appropriately, whether the sample size was adequate, whether the condition was measured with a valid and reliable method, and whether the response rate was high enough to trust. A study with a convenience sample and a low response rate can report a precise-looking proportion that is badly unrepresentative, and only a structured appraisal catches it. The appraisal then feeds your assessment of certainty in the evidence and your decision about which studies belong in the pooled estimate at all.
Extracting the numerator and denominator
The single most important extraction rule in a prevalence meta-analysis is to record each study’s numerator (the number of cases) and denominator (the number sampled), not merely the reported percentage. The counts are what let the software apply the correct variance-stabilising transformation and the exact binomial variance, and they are essential if you later fit a generalised linear mixed model that works directly on the counts. Percentages alone force you to reconstruct the counts by back-calculation, which fails when a study rounds heavily or reports a proportion without its sample size. Where a study gives a proportion and a confidence interval but no counts, you can sometimes recover an effective sample size from the interval width, but this is a workaround to flag as a limitation, not a first choice. Extracting counts also makes your effect data auditable, so a reader can reproduce the pooled figure from your table. Record alongside each count the year, setting, age range, and case definition, because those fields become the moderators you will later test to explain the heterogeneity, and going back to re-extract them once screening is finished wastes hours that a well-designed extraction form would have saved.
Presenting the pooled estimate and the spread
How you display the result shapes how honestly it is read. A forest plot of the individual study proportions with the pooled estimate at the foot is the standard, and it should show the study weights so a reader sees whether one large survey is carrying the estimate. Add the prediction interval as a distinct band beneath the pooled diamond, because in a prevalence review it is usually much wider than the confidence interval and communicates the real uncertainty about a new setting. Report the pooled proportion as a percentage with its confidence interval, the prediction interval, the number of studies and total participants, the I-squared, and the transformation used. Where you present subgroups, give each subgroup its own pooled estimate rather than only a test of interaction, since readers want the prevalence in their population of interest, not just a p-value for whether groups differ.
Subgroups and moderators that usually matter
Because heterogeneity in prevalence work is driven by real differences, the moderators you test are rarely arbitrary. Geographic region or country is almost always worth examining, because prevalence varies with environment, health systems, and genetics. The diagnostic criteria or case definition matters enormously, since a broader definition mechanically raises measured prevalence, and mixing definitions without a subgroup is a classic error. Age band and sex shape most conditions, and study year can reveal secular trends. Setting, whether community, primary care, or hospital, changes the sampled population and therefore the estimate. Testing these through a meta-regression of prevalence moderators does two things at once: it may explain a chunk of the heterogeneity, and it produces the population-specific estimates that make the review clinically useful. Pre-specify the moderators in the protocol so the analysis is confirmatory rather than a hunt for a significant subgroup.
A practical workflow and where to start
A defensible prevalence meta-analysis moves through a clear sequence: extract each study’s numerator and denominator (not just the reported percentage), appraise every study with the JBI checklist, choose and pre-specify a transformation, pool under random effects, report the estimate with a confidence interval and a prediction interval, and explain the heterogeneity through subgroups or meta-regression. Record all of this against the PRISMA 2020 checklist so the choices are transparent. You can sanity-check a single study’s proportion and its confidence interval in our single proportion calculator before committing to the pooled model, and view the combined estimate in the meta-analysis calculator. The recurring mistakes are three: pooling raw proportions and letting a near-zero study dominate, using the double arcsine back-transformation uncritically after the 2019 warning, and reporting a single pooled percentage as if the vast heterogeneity did not exist. Avoiding those three is most of the battle.