A realist review, also called a realist synthesis, is a theory-driven approach to evidence synthesis that asks not whether an intervention works but how, why, for whom, and in what circumstances it works. Developed by Ray Pawson, it builds and tests a programme theory by tracing context-mechanism-outcome configurations across a deliberately varied body of evidence, and it reports against the RAMESES publication standards rather than the conventions of an effectiveness review.

The engine of the method: context, mechanism, outcome

Realist logic starts from a simple observation: interventions do not work by themselves. A smoking cessation programme, a hospital checklist, or a school feeding scheme is only ever a set of resources and opportunities; what produces change is how people reason and respond to those resources, and that response depends on the setting. The unit of analysis is therefore the context-mechanism-outcome configuration, usually written as C plus M leads to O. The context is the circumstance that makes a response possible: staffing levels, organisational culture, trust in authority, poverty, prior experience. The mechanism is the underlying reasoning or reaction the intervention triggers in that context: fear of sanction, restored confidence, peer pressure, a sense of being monitored. The outcome is what follows, intended or not. The same intervention can fire different mechanisms in different contexts and so produce different outcomes, which is precisely why effect estimates for complex interventions swing so wildly between trials. A realist review treats that variation not as statistical noise to be averaged away but as the primary evidence to be explained.

A concrete example makes the logic visible. Suppose a review of text-message reminders for medication adherence finds trials ranging from strong effects to none. A realist reading might propose that in contexts where patients already trust their clinic, the reminder fires a mechanism of feeling cared for, and adherence rises; in contexts where patients fear surveillance or stigma, the same message fires a mechanism of feeling monitored, and some patients disengage entirely. Two configurations, one intervention, opposite outcomes. Once stated, each configuration becomes a testable proposition the review can pursue across the literature, and the scattered trial results stop being a contradiction and start being data.

How Pawson’s approach differs from an effectiveness review

A conventional systematic review of effectiveness asks a verdict question: does the intervention work, on average, compared with a control? It answers by exhaustive searching, duplicate screening, checklist-based appraisal, and, where possible, pooling. Pawson’s argument was that for complex social interventions this machinery answers the wrong question, because an average effect across incompatible contexts tells a policymaker almost nothing about what will happen in their setting. A realist review instead asks an explanatory question and organises every methodological choice around it. Searching is purposive rather than exhaustive. Any evidence that can speak to the theory is admissible, including process evaluations, qualitative studies, and searching grey literature sources for policy documents and evaluation reports. There is no hierarchy of evidence with randomised trials at the top, because a small qualitative study may explain a mechanism that a large trial can only measure. The output is not a pooled estimate but a refined, tested theory of how the intervention generates its effects. It is a systematic, auditable process, which separates it sharply from an unstructured narrative review, but its logic of inquiry is interpretive and explanatory throughout. Where it sits among the wider family of literature review types is closest to the theory-building end of the spectrum.

Programme theory: the spine of the review

Every realist review begins by drafting an initial programme theory: a candidate explanation of how the intervention is supposed to work, for whom, and under what conditions. The draft is assembled from the intervention’s own stated logic, background literature, formal theories from psychology or sociology, and interviews with programme architects and practitioners. It is deliberately provisional. The rest of the review exists to test and refine that theory against empirical evidence: each included document is interrogated for what it says about the hypothesised contexts, mechanisms, and outcomes, and the theory is progressively adjusted, elaborated, or overturned. Middle-range theories borrowed from other disciplines often do the heavy lifting here; a review of audit and feedback programmes, for instance, might recruit control theory or social comparison theory to explain why feedback changes some clinicians’ behaviour and not others’. By the end, the review presents a refined programme theory, often expressed as a set of demi-regularities, the semi-predictable patterns of how the intervention behaves across contexts. This gives decision-makers something an effect estimate cannot: a transferable explanation they can map onto their own setting before committing resources.

Iterative, purposive searching

Searching in a realist review is organised around the theory, not around a fixed protocol executed once. A background search scopes the territory and feeds the initial programme theory. A broad main search then assembles the core body of evidence. As the theory develops, focused purposive searches follow: hunting for evidence on one specific mechanism, borrowing a formal theory from another discipline, or deliberately seeking cases where the intervention failed, because failures are where mechanisms show themselves most clearly. Techniques such as citation tracking, snowballing, and cluster searching for sibling papers around a key evaluation matter more here than exhaustive database coverage. Searching stops at theoretical saturation, the point where new documents stop changing the theory, rather than when a record count is exhausted. Every search is still documented and auditable; iteration is a feature of the design, not a loosening of standards, and it is one reason the method should not be confused with the pragmatic streamlining of rapid review methods.

Appraisal by relevance and rigour, not checklists

Realist reviews do not appraise studies with design-specific checklists, because whole-study quality scores answer the wrong question. A methodologically flawless trial may contribute nothing to the theory, while a modest process evaluation may contain the single observation that unlocks a mechanism. Each source is instead judged on two questions. Relevance: can this document speak to the programme theory under test? Rigour: is the specific claim being borrowed from it credible and trustworthy, given how the underlying work was done? Appraisal therefore happens at the level of the individual data fragment, not the whole study, and the reviewer’s judgements are recorded so the chain from evidence to theory stays auditable. The analytic work that follows, tagging fragments of text against candidate contexts, mechanisms, and outcomes, closely resembles coding data in qualitative reviews, and teams with qualitative synthesis experience adapt to it fastest. A practical safeguard is to record, for every configuration in the final theory, which documents support it, which contradict it, and how the contradiction was resolved. That evidence table is what reviewers of the finished synthesis will ask for first, and building it as you go is far cheaper than reconstructing it at the end.

Reporting with the RAMESES standards

The RAMESES publication standards (Realist And Meta-narrative Evidence Syntheses: Evolving Standards) are the recognised reporting framework, playing the role PRISMA plays for conventional reviews. The realist list runs to nineteen items covering the rationale for a realist approach, how the initial programme theory was built, how searching evolved, how relevance and rigour were judged, and how the refined theory was produced from the data. RAMESES also publishes quality standards that distinguish adequate from exemplary conduct, and journal editors and funders increasingly expect both to be cited and followed. Two practical implications follow. First, document decisions as you make them, because the standards require you to narrate the review’s intellectual journey, including theories tried and abandoned. Second, publish a protocol early even though the method is iterative: the protocol fixes the question and the approach while explicitly flagging the elements that are expected to evolve. Reviewers familiar with conventional reporting sometimes mistake that declared flexibility for weakness, and a protocol that anticipates the objection, explaining which decisions are fixed and which are theory-driven, saves an awkward exchange at peer review.

Realist review or realist evaluation

The two labels are siblings, not synonyms, and confusing them is the most common terminology error in grant applications. A realist review is secondary research: it builds and tests programme theory using evidence that already exists in published and unpublished documents. A realist evaluation is primary research: it tests programme theory by collecting new data, typically interviews, observations, and outcome measures, inside a live programme. The two are frequently sequenced, with a review producing the candidate configurations that an evaluation then tests in the field, and RAMESES publishes separate standards for each. Funders increasingly commission the pair together, because a review alone cannot observe mechanisms directly and an evaluation alone wastes effort rediscovering what the literature already shows.

When a realist review is the right choice

The method earns its cost when three conditions line up. The intervention is complex: multiple interacting components, dependent on human behaviour, sensitive to setting, such as integrated care models, community health worker programmes, pay-for-performance schemes, or school-based interventions. The existing evidence is mixed or contradictory, with trials pointing in different directions and nobody able to say why. And the decision-makers need implementation guidance, not a verdict: they already intend to act and want to know how to make the intervention work in their context. If instead the question is a clean effectiveness comparison of a well-defined intervention, a conventional review remains the right tool, and if the field is already saturated with systematic reviews, an umbrella review of existing reviews may serve better. Realist reviews also pair naturally with qualitative approaches, and our qualitative evidence synthesis service covers the neighbouring designs.

Team and timeline

A credible realist review is a team sport. The typical core is a lead reviewer with realist training, a second reviewer for coding and theory-testing discussions, an information specialist comfortable with iterative searching, and a content expert who knows the intervention’s world; many teams add an experienced realist methodologist as adviser, and the RAMESES community actively supports newcomers. Expect nine to eighteen months for a full synthesis, with the theory-building and theory-testing phases, not the searching, consuming most of it. A rough shape: two to three months to scope the territory and draft the initial theory, three to six months of iterative searching, appraisal, and coding, three to six months of theory refinement and stakeholder testing, and the remainder for writing against the reporting standards. Compressed versions exist for urgent policy questions, but the compression must be declared, and the intellectual heart of the method, the disciplined refinement of a programme theory against evidence, cannot be skipped without producing something that is realist in name only.