Deduplicating search results is the step that removes records returned by more than one database so that each unique study is screened only once. Because a systematic review searches several databases that index overlapping journals, the combined export routinely contains a large share of duplicate records, and removing them before screening is both a courtesy to your reviewers and a requirement for an honest flow diagram.
Why de-duplication has to happen before screening
Two databases will often return the same article under slightly different metadata: a different page range, an abbreviated journal title, a missing digital object identifier. If those copies reach your reviewers, the same study is screened twice, wasting effort and risking contradictory decisions on what is really one record. Removing duplicates first means every screening decision is made once, and the count of unique records is the figure that anchors the top of your PRISMA flow diagram.
What counts as a duplicate
Exact and near matches
Some duplicates are trivial: identical digital object identifiers or identical titles and authors. The hard cases are near duplicates, where punctuation, accented characters, or a truncated title stop an automated tool from matching two records that a human can see are the same paper. A robust process therefore combines automated matching on stable fields with a human review of the borderline pairs.
Distinguish duplicates from companion reports
A genuine duplicate is the same report retrieved twice. It is different from multiple reports of one study, such as a trial described in both a protocol paper and a results paper. Those are not deleted; they are linked so that the study is counted once while every report informs extraction. Confusing the two leads either to lost data or to double-counting a single study in your synthesis.
Which fields to match on, and in what order
Reliable de-duplication is a cascade from the most stable field to the least. The digital object identifier is the strongest anchor: two records sharing one are almost always the same paper. Next comes a normalised combination of title plus first author plus year, then title plus journal plus volume and pages. The weakest signal is title similarity alone, which is where near matches hide. A sensible matching order runs through these in turn:
- Remove records that share an identical digital object identifier, which are safe to merge automatically.
- Remove records with identical normalised titles and matching first author and year.
- Flag, but do not auto-delete, records whose titles are highly similar but differ in punctuation, accented characters, or truncation.
- Have a human confirm each flagged pair, since this is where a tool either misses a true duplicate or wrongly merges two distinct papers.
How to de-duplicate reliably
Reference managers such as EndNote and Zotero and purpose-built screening platforms both offer automated de-duplication, and the two complement each other. The practical workflow is to export from every database in a consistent format, import everything into one library, run automated removal on exact identifiers, then manually scan the remaining near matches. Many teams do the final pass inside a dedicated systematic review screening tool, which records the number removed so the figure flows straight into reporting. A defensible process is conservative on the borderline pairs: when in doubt, keep both records and let the first screening pass catch the survivor, because wrongly deleting a unique study at this stage is invisible and unrecoverable in the same way a wrongful exclusion is. Standardised reporting frameworks now exist for this step, the most cited being the four-step method that records counts before and after each automated and manual pass.
Why exported metadata disagrees across databases
Understanding why the same article looks different in two exports makes the manual pass far quicker. Databases apply their own house style to author names, so one indexes “Smith J” and another “Smith John A”. Journal titles are abbreviated by some sources and spelled out by others, so “J Clin Epidemiol” and “Journal of Clinical Epidemiology” describe the same venue. Accented characters survive in one export and are stripped in another, and page ranges differ where one source records an electronic article number rather than printed pages. None of these are errors; they are cataloguing conventions, and a reviewer who recognises them stops second guessing whether two near-identical records are truly the same paper. The practical implication is that matching on a normalised version of each field, with case, punctuation, and accents folded away, catches far more true duplicates than matching on the raw text the database happened to supply.
Recording the numbers for reporting
The PRISMA 2020 flow diagram asks for two specific de-duplication figures: the number of records identified across all sources, and the number of duplicates removed before screening. Those numbers must reconcile, so capture them at the moment you de-duplicate rather than reconstructing them later. Keeping the export files and the de-duplicated library is part of documenting and reporting your search, and it lets anyone audit how you arrived at the set that entered title and abstract screening.
Common de-duplication mistakes
A handful of errors recur. The first is trusting an automated tool blindly and never inspecting the borderline pairs, which silently loses unique studies whose metadata happened to look similar. The second is the opposite, over-merging, where two genuinely different papers with near-identical titles are collapsed into one. The third is failing to distinguish a duplicate from a companion report, so a trial described across a protocol paper and a results paper is either wrongly deleted or wrongly counted twice. The fourth is de-duplicating after screening has started, by which point reviewers may already have cast contradictory votes on the same study. Each of these distorts the figures that anchor the PRISMA 2020 reporting standard, so a clean process catches them before a single human reads a record.
Where de-duplication sits in the pipeline
De-duplication is the bridge between the search and the screen. It follows your systematic literature search across databases and grey literature, and it hands a single clean record set to your two reviewers. The same overlap problem returns whenever new records arrive mid-review, for example through backward and forward citation searching, so any late additions are de-duplicated against the existing library before they enter screening. Done well, de-duplication is invisible; done poorly, it either inflates your workload with repeated records or, worse, lets the same study be counted twice in the final synthesis, where a single trial entered twice silently doubles its weight in the pooled estimate.