A living systematic review is a systematic review that is continually updated, with new evidence incorporated as soon as it becomes available. Instead of being finished, published, and left to age, the review is underpinned by active surveillance: searches re-run on a fixed schedule, often monthly, with eligible new studies screened, appraised, and folded into the synthesis, and the published version revised whenever the evidence changes enough to matter. The method exists because in fast-moving fields a conventional review can be out of date before it appears, while a living review keeps the answer current for as long as the question stays important.

Why the conventional review model goes stale

A conventional review is a snapshot. The team works through the standard steps of a systematic review, the search closes months before publication, and from that day the review begins to decay. Empirical work on review currency has found that a meaningful fraction of reviews have conclusions overturned or substantially modified within a few years, and in rapidly evolving areas the half-life is far shorter. The traditional remedy, a full update every few years, is slow, expensive, and repeats much of the original labour, and many reviews are simply never updated at all. Meanwhile decision makers keep citing the published version, unaware that the evidence has moved. The living model reframes the problem: rather than treating an update as a new project, it treats currency as a property of the review itself, bought with a continuous, smaller stream of effort instead of periodic heroic efforts.

The Cochrane living evidence framework

Cochrane piloted living systematic reviews from 2017 and has since built the most developed framework for them. Its guidance defines living mode as a decision about how a review is maintained, not a different review type: the eligibility criteria, appraisal standards, and synthesis methods are those of any rigorous review, with three additions. First, a published living protocol states the surveillance schedule, the update triggers, and how long living mode will run. Second, search and screening become recurring processes with named owners rather than one-off tasks. Third, the review carries explicit version and currency statements so readers can see when the evidence was last checked, even when nothing changed. Cochrane’s framework also stresses transition planning: living mode should end when the question stabilises, and the guidance treats retiring a living review as a legitimate, planned outcome rather than a failure.

The framework has spread well beyond Cochrane. Guideline agencies, health technology assessment bodies, and several journals now publish or accept living formats, and the same logic has been extended to living network meta-analyses that keep whole treatment hierarchies current rather than a single comparison. What all the implementations share is the insistence that “living” describes an operational commitment with named owners and a schedule, not a label added to a review the authors merely hope to update.

When living mode is justified

Living mode is demanding, so the framework reserves it for questions that pass three tests. The question must be a priority for decision making, typically because a guideline, policy, or active clinical controversy depends on it. The evidence must be moving, with new trials expected at a rate that could plausibly change the answer; a field producing two relevant trials a decade does not need monthly surveillance. And the existing evidence must leave meaningful uncertainty, because once an answer is precise and stable, further updating adds nothing. A useful thought experiment is to ask what a new trial would have to show to change the conclusion, a question closely related to how certainty of evidence is rated: when certainty is already high, the trigger threshold is rarely reachable and living mode is wasted effort.

The classic candidates are new drug classes with a heavy trial pipeline, emerging infectious diseases, rapidly evolving digital and surgical interventions, and any question feeding a living guideline. Poor candidates are settled questions, historical exposures, and fields where the plausible flow of new studies is a trickle. The three criteria should be revisited on a schedule, because a question that justified living mode in year one may fail the test by year three, at which point the review should transition back to conventional maintenance with its final version clearly stamped.

Continual search surveillance

The engine of a living review is the surveillance search. The baseline strategy is designed once, across the core databases for systematic reviews, and then re-executed on a schedule, with monthly the most common cadence. Teams supplement database re-runs with automated alerts and push services, trial registry monitoring for newly completed studies, and preprint server checks in fields where results appear there first. Each surveillance cycle produces a small batch of records that flows into the standing screening process, and the review logs the date and yield of every cycle. That log is itself a finding: a run of empty cycles is evidence that the field has gone quiet and living mode may be ready to retire.

Surveillance design involves a real trade-off between sensitivity and sustainability. A maximally sensitive strategy re-run monthly can bury a small team in records, so living teams often maintain two tiers: a pragmatic, precise strategy for the monthly cycle, and a periodic full-sensitivity sweep, perhaps annually, to catch anything the lean strategy missed. Registry surveillance deserves equal weight, because watching a registered trial move to “completed” status months before publication lets the team anticipate an update rather than react to it, and approach the investigators early for results.

Update triggers and frequency

Searching continuously does not mean republishing continuously. Living reviews separate surveillance frequency from update frequency using pre-specified update triggers. Minor triggers, such as a small study consistent with the existing pooled effect, may only update the study tables and the currency statement. Major triggers force a full re-synthesis: a large trial, a change in the direction or precision of the pooled estimate, new evidence on harms, or a study that alters the certainty rating. Stating the triggers in the protocol matters because it protects the team from both failure modes, neglect on one side and churn on the other, and it tells readers exactly what “up to date” means for this review. In practice reported cadences vary widely, and studies of published living reviews find many drift from their stated schedule, which is precisely why an explicit, resourced schedule beats good intentions.

Setting up a living review step by step

Standing a living review up follows a recognisable sequence. It starts with a rigorous baseline review, because living mode maintains quality rather than creating it; a weak review updated monthly is just a weak review with a timestamp. The protocol is then extended with the living elements: the surveillance cadence, the update triggers, the roles responsible for each cycle, and the planned duration of living mode. Next the team engineers the pipeline so each cycle is cheap: saved search strategies that re-run automatically, deduplication against the master library, a trained screening classifier, extraction forms that append to a structured dataset, and analysis code that regenerates every figure and table from that dataset on demand. The first few cycles are treated as a shakedown, timing each stage and fixing the bottlenecks, before the rhythm settles into a predictable monthly routine. Teams writing their first review of any kind should master the craft of writing a systematic review before attempting to keep one alive, because living mode amplifies whatever discipline, or indiscipline, the baseline process contains.

Workflow and technology

A living review is an operations problem as much as a methods problem, and technology carries much of the load. Machine learning screening assistance, of the kind built into modern screening tools for systematic reviews, is close to essential: classifiers trained on the review’s existing decisions rank each surveillance batch so humans check the likely includes first, and some platforms auto-exclude records below a validated threshold. Automated search re-execution, deduplication against the master library, living flow diagrams, and data pipelines that regenerate the meta-analysis from a structured dataset rather than hand-edited tables all shrink the marginal cost of a cycle. The human workflow changes too: instead of a project team that disbands, a living review needs a small standing rota with named responsibility for each monthly cycle, which is a staffing commitment as real as any technical one.

Living guidelines and versioned publication

The main consumers of living reviews are living guidelines, guideline programmes that revise recommendations as their underlying reviews update. Australia’s living stroke guidelines and the World Health Organization’s living COVID-19 recommendations both run on this pipeline, and the connection is symbiotic: the guideline supplies the justification for surveillance, and the review supplies the currency the guideline advertises. Publication has had to adapt as well. Journals and platforms now support versioned publication, where each update appears as a citable version with a shared identifier and a visible change history, and Cochrane treats the living review as a single record whose citation carries a version date. Reporting each version still follows the PRISMA 2020 reporting guideline, with flow diagrams and study counts updated per cycle so the audit trail survives across versions.

The COVID-era track record, and what a living review honestly costs

The COVID-19 pandemic was the method’s proving ground and its stress test. Living reviews and living network meta-analyses of competing treatments delivered genuinely current answers on therapies at a speed conventional reviewing could never have matched, and directly fed the living guidelines that steered treatment worldwide. The honest half of the record is that many living reviews launched in that period lapsed: audits of pandemic-era living reviews found that a substantial share were never updated after their first version, and only a minority sustained a regular cadence. The lesson is not that the model fails but that it costs what it costs. Budget for the baseline review plus a recurring commitment, roughly a fifth to a half of the original effort per year depending on the field’s pace, secure that funding before promising a living review, and plan the exit. A review that lives for two well-maintained years and retires deliberately serves readers far better than one that promises immortality and silently stops.

Where does the recurring cost actually go? Screening the monthly yield is usually modest once a classifier is trained; the expensive triggers are the cycles where a major study lands and forces re-extraction, re-appraisal, re-synthesis, and a fresh round of editorial and peer review for the new version. Information specialist time for maintaining search strategies, statistician time for re-running analyses, and coordination overhead for the standing rota make up the rest. Teams that thrive treat these as budgeted line items with named owners; teams that fold treated them as volunteer goodwill. That, more than any methodological subtlety, is what separates the living reviews still current today from the ones that survive only as an optimistic first version.