Inter-rater reliability in screening is a measure of how consistently two reviewers reach the same include-or-exclude decision on the same records, after correcting for the agreement you would expect by chance alone. The most common statistic is Cohen’s kappa, and reviewers report it from a pilot round to show that their eligibility criteria are being applied the same way by both screeners.

Why raw agreement is misleading

If you simply count the proportion of records two reviewers agreed on, the figure looks impressive but is deceptive. When most records are clearly irrelevant, two reviewers who both reject almost everything will agree on the vast majority of records without ever testing their judgement on the hard cases. Percent agreement rewards this. A chance-corrected statistic strips out the agreement you would expect from two people guessing at the same base rate, leaving only the agreement that reflects genuinely shared criteria. That is why journals ask for kappa rather than a raw percentage.

Cohen’s kappa and how to read it

What the value means

Cohen’s kappa ranges from below zero to one. A value of one is perfect agreement, zero is no better than chance, and common interpretation bands treat values above roughly 0.6 as substantial and above 0.8 as almost perfect. These thresholds are conventions, not hard laws, but they give you a yardstick: if your pilot kappa is low, your criteria are not yet operational. You can compute the statistic from your two reviewers’ votes with our Cohen’s kappa calculator, which takes the agreement table and returns the corrected figure.

The interpretation bands in practice

The most widely cited yardstick is the Landis and Koch scale, which labels kappa below 0.00 as poor, 0.00 to 0.20 as slight, 0.21 to 0.40 as fair, 0.41 to 0.60 as moderate, 0.61 to 0.80 as substantial, and 0.81 to 1.00 as almost perfect. For screening, most teams treat a pilot kappa at or above 0.60 as the threshold to begin the full screen, and a value below roughly 0.40 as a clear instruction to stop and rewrite the rules. Worked through a concrete pilot, the logic is intuitive: if two reviewers each screen 80 abstracts and agree on 72, raw agreement is 90 per cent, but if eligible records are rare the chance-corrected kappa may sit near 0.45, which tells you the apparent agreement is mostly two people rejecting obvious irrelevance in lockstep rather than sharing a sharp definition of inclusion.

The paradox to watch for

Kappa has a known quirk: when one category is very rare, as eligible studies usually are, kappa can be low even when percent agreement is high. This is the kappa paradox. It does not mean your screening is poor; it means the prevalence of inclusions is so skewed that the statistic behaves oddly. Report kappa alongside the raw counts so a reader can see the full picture rather than a single number out of context. A practical safeguard is to report three figures together: the observed agreement, the kappa, and the number of records both reviewers marked for inclusion. When inclusions are very sparse, a prevalence-adjusted statistic or simply the proportion of positive agreement can be more honest than kappa alone, and saying so in the methods pre-empts a reviewer querying a deceptively low figure.

When and how to measure it

Reliability is measured during the pilot, before the full screen, on a shared sample of 50 to 100 records that both reviewers assess independently. A weak pilot kappa is a gift: it tells you to fix your inclusion and exclusion criteria before you screen thousands of records rather than after. Some teams re-check reliability partway through a long screen to catch drift, where reviewers gradually diverge as fatigue sets in. The same logic applies when you decide how many reviewers screening needs: the second reviewer only adds value if both are calibrated to the same rules.

A step-by-step way to run the pilot

Treat the reliability check as a small, repeatable procedure rather than a one-off statistic, and the criteria almost always sharpen themselves:

  1. Draw a random sample of 50 to 100 records from the de-duplicated set, ideally enriched with a few you already suspect are eligible so the rare category is represented.
  2. Have both reviewers screen the sample independently, recording an include or exclude vote for every record with no discussion.
  3. Build the two-by-two agreement table and compute kappa, for example with our chance-corrected agreement calculator.
  4. Pull out every disagreement and ask, for each, whether the rule or the reviewer caused it; a cluster of conflicts on one criterion means the rule is ambiguous.
  5. Rewrite the offending criterion in operational language, then re-pilot on a fresh sample until kappa clears your pre-agreed threshold.

This loop is the cheapest insurance in the whole review, because every hour spent calibrating two reviewers on 80 records saves days of re-screening across several thousand. It also produces a defensible sentence for the methods: the team can state the pilot sample size, the achieved kappa, and the criteria revisions that followed, all of which a peer reviewer reads as evidence of a disciplined first screening pass.

Reliability beyond two raters

Cohen’s kappa is built for exactly two reviewers. When three or more screen the same records, a related statistic such as Fleiss’ kappa generalises the idea to multiple raters. The principle is identical: correct for chance, then judge whether the team is applying one shared standard. The same family of agreement statistics reappears later in the review, for instance when two assessors compare judgements during risk of bias assessment. A further variant, weighted kappa, is reserved for ordered categories, so it has little place in a binary include-or-exclude screen but reappears when reviewers rate something on a graded scale. The point to retain is that the choice of statistic follows the structure of the decision: two raters and two categories call for Cohen’s kappa, three or more raters call for Fleiss’ kappa, and ordered ratings call for the weighted form.

Common mistakes when reporting reliability

A few errors recur often enough that reviewers learn to look for them. The first is reporting only percent agreement and calling it reliability, which ignores chance entirely. The second is computing kappa on the whole screen after the fact, when its real value is as a pre-screen diagnostic that can still change the criteria. The third is treating a single threshold as a pass or fail gate without inspecting the disagreements themselves, since the qualitative pattern of conflicts is more informative than the number. A fourth is calculating reliability after the reviewers have already discussed the records, which inflates the figure and defeats the purpose, because the statistic must reflect independent judgement made blind. Avoiding these keeps the figure meaningful and stops it becoming a box-ticking ritual disconnected from the quality of the eligibility rules it is meant to test.

What reliability does and does not prove

A high kappa proves your two reviewers are consistent; it does not prove they are correct. Two reviewers can agree perfectly while both applying a criterion that misreads the review question. Reliability is necessary but not sufficient: it confirms the screen is reproducible, which then has to sit on top of well-framed criteria drawn from a sound research question. Reported together, a clear question, sharp criteria, and a strong kappa are what let readers trust the studies you selected.