An evidence gap map is a systematic, visual overview of what research exists on a topic, built as a matrix of interventions against outcomes. Each row is an intervention, each column is an outcome, and each cell shows how many studies have examined that particular combination, usually as a bubble sized by the volume of evidence. Full cells reveal where the evidence is concentrated; empty cells reveal the evidence gaps. The method searches and codes studies with full systematic rigour but deliberately stops short of synthesising their results, because its job is to show where the evidence sits, not what it says.

The 3ie and Campbell tradition behind the method

Evidence and gap maps grew out of international development evaluation. The International Initiative for Impact Evaluation, known as 3ie, developed the format in the early 2010s to help funders see at a glance which development programmes had been rigorously evaluated and which had not, and it now hosts a large public collection of maps on themes from agriculture to social protection. The Campbell Collaboration then formalised the approach for the social sciences, publishing dedicated methodological guidance and peer-reviewed maps under the label evidence and gap map, the term most methodologists now prefer. That lineage matters because it fixed the defining features of the genre: a stakeholder-driven framework, a systematic and documented search, transparent coding, and a commitment to display rather than synthesis. A map produced to this standard is a piece of research infrastructure in its own right, not a lighter version of a review.

Reading the matrix: rows, columns, bubbles, and filters

The heart of any map is the framework matrix. Rows carry the intervention taxonomy, often grouped into broader categories such as prevention, treatment, and service delivery. Columns carry the outcome taxonomy, running from intermediate outcomes through to final impacts. A study lands in every cell its comparisons touch, so one trial can appear in several cells. Most maps encode extra information in the display itself: bubble size shows the volume of studies, colour distinguishes study designs or the confidence rating of included systematic reviews, and interactive filters let a user cut the matrix by population, region, setting, or study design. The result is a dense but legible picture. A policymaker can see within seconds that, say, school feeding programmes have dozens of impact evaluations on attendance but almost nothing on long-term earnings, which is precisely the kind of research prioritisation signal the format was designed to deliver.

Not every empty cell is a genuine gap, and reading a map well means knowing the difference. Some combinations are empty because they are implausible or unethical to study, and a good report says so rather than inviting funders to fill them. Methodologists also distinguish an absolute gap, a cell with no primary studies at all, from a synthesis gap, a cell with a healthy cluster of primary studies but no trustworthy review pulling them together. The two call for opposite responses: an absolute gap needs new primary research, while a synthesis gap needs a review, which is faster and cheaper to commission. Mature maps make this distinction visible by plotting primary studies and existing reviews as separate segments within each bubble, so the reader can diagnose the cell at a glance instead of merely counting.

How a map differs from a scoping review and a systematic review

The three formats answer different questions. A systematic review asks “what does the evidence say about this focused question?” and synthesises results, often statistically. A scoping review asks “what is known about this broad topic?” and reports a narrative or tabular description; the practical mechanics are covered in our guide to conducting a scoping review. An evidence and gap map asks “where does the evidence sit, and where is it missing?” and answers with a visual database rather than prose. Two consequences follow. First, a map is organised by a pre-specified framework rather than by themes that emerge during the work, which makes it more structured than most scoping exercises; the boundary between the formats is discussed further in how systematic and scoping reviews differ. Second, a map makes no attempt to synthesise effects. It will tell you that eleven trials examined an intervention and outcome pair, but not whether the intervention works. Maps therefore sit alongside umbrella reviews and mapping reviews in the wider family of literature review types as a “big picture” format that often precedes, and scopes, a later synthesis.

Developing the framework before touching the literature

The framework is the intellectual core of the project and it must be settled before searching begins. Teams draft the intervention and outcome taxonomies from existing theory, programme logic models, and prior reviews, then refine them with stakeholder consultation: funders, practitioners, and people affected by the interventions all get a say in which rows and columns matter. The framework then behaves like a protocol. Every included study must be codeable against it, so vague categories, overlapping outcomes, or a missing row will surface as friction during coding and force expensive rework. Good practice is to pilot the framework on twenty or thirty known studies, adjust, and lock it, recording any later amendments openly. A well-built framework also defines the map’s dimensions and filters, the extra variables such as region, population, and study design that make the interactive version genuinely useful rather than a static picture.

Systematic searching and coding, without synthesis

From the framework onwards, the workflow borrows its rigour wholesale from systematic reviewing. Searches are designed and documented across the bibliographic databases used for systematic reviews, supplemented by grey literature sources, specialist repositories, and citation chasing, because gaps are only meaningful if the search that failed to fill them was genuinely comprehensive. Records are screened against explicit eligibility criteria, ideally in duplicate and often with help from machine-assisted screening software given the volumes involved. Included studies are then coded, not extracted for synthesis: reviewers record which framework cells each study occupies, its design, population, region, and other filter variables. Many maps also include existing systematic reviews as a distinct evidence type and appraise them with the AMSTAR 2 tool so users can see at a glance whether a cell already contains a trustworthy synthesis. What never happens is pooling. Effect sizes are not combined, and the map draws no conclusions about effectiveness, which is exactly what keeps it quick to update and neutral for priority-setting.

EPPI-Mapper and the interactive map itself

The standard delivery vehicle is EPPI-Mapper, a free application from the EPPI Centre at University College London. Teams manage screening and coding in EPPI-Reviewer or a comparable platform, export the coded dataset, and EPPI-Mapper turns it into a self-contained interactive web page: a zoomable matrix with clickable bubbles, segment colouring by design or confidence rating, and filter panels for every coded dimension. Because the output is a single file, it can be hosted anywhere, embedded in a funder’s website, and versioned as the evidence base grows. Some teams go further and run the map as a living evidence and gap map, re-running searches on a schedule and republishing, an approach that borrows its surveillance logic from the living systematic review model. Either way, the interactive format is not decoration. The ability to filter to “randomised studies, low-income countries, adolescents” and watch the gaps move is what turns a static figure into a decision tool.

How funders use maps to set research priorities

For research funders and commissioning bodies, the map is chiefly an investment planning instrument. Crowded cells tell a funder that a new primary study may add little, and that money is better spent commissioning a synthesis of what already exists, sometimes a full systematic review, sometimes a review of existing reviews where several syntheses already compete. Empty or sparse cells with high policy salience become the research agenda: several development funders now require grant applicants to situate proposals against a published map, and organisations such as the World Health Organization have used maps to structure entire research programmes. Maps also expose synthesis gaps, cells with plenty of primary studies but no recent trustworthy review, which are the cheapest wins in the whole evidence system. The honest caveat is that a map records quantity, not quality of answer, so a full cell can still hide weak or conflicting evidence, and funders read bubble counts alongside the design and confidence colouring for that reason.

Common pitfalls that undermine a map

The failures of weak maps are predictable. The most common is framework drift: categories added or merged midway through coding without documentation, which silently changes what the bubbles mean and makes the finished matrix impossible to audit. The second is treating the map as a synthesis, sliding from “eleven studies exist here” to “this intervention works”, a claim the method cannot support and one that careful reports rule out explicitly in their limitations. Third comes screening on a single reviewer to save time; with tens of thousands of records the error rate compounds, and an unreliable map is worse than none because its gaps may simply be studies the screeners missed. Fourth, skipping the appraisal of included systematic reviews leaves users unable to tell whether a promising cell already contains a reliable answer or a clutch of contradictory, low-quality syntheses. Finally, a map that is never refreshed decays quickly in active fields, so commissioners should decide at the outset whether the budget covers periodic re-searches or whether the map is explicitly dated as a one-off snapshot. None of these pitfalls is subtle, which is precisely why funders scrutinise the methods section of a map as closely as they would a review’s.

Typical scale, team, and timeline

Maps are large undertakings. Because the questions are broad, searches commonly return ten to fifty thousand records, and finished maps frequently include several hundred to a few thousand studies. A realistic team pairs an information specialist with two or more coders and a methodologist who owns the framework. Campbell-standard maps typically take six to twelve months from protocol to published map, with screening and coding consuming most of that time; a narrowly framed map on a contained topic can be delivered in three to four months, while a very broad sector map with stakeholder workshops can run longer. The economics still compare well with synthesis, because coding a study for a map is far faster than extracting and appraising it for pooling. Teams planning a map as the first stage of a larger programme, with focused syntheses to follow in the cells that justify them, should design the coding so it can be reused downstream, a sequencing pattern we cover across the standard stages of a systematic review. Done that way, the map pays for itself twice: once as a published decision tool, and again as the screened, coded foundation for every review that follows.