A CASP checklist is a structured appraisal instrument from the Critical Appraisal Skills Programme, a United Kingdom initiative that publishes free, design-specific question sets for judging published research. Each checklist walks an appraiser through roughly ten to twelve questions organised into three sections, asking whether the results of a study are valid, what the results actually are, and whether they will help locally, without ever producing a numeric score.

Where the checklists came from and when to reach for them

The Critical Appraisal Skills Programme grew out of the evidence-based-medicine movement in Oxford in the early 1990s, where it was built to teach clinicians, commissioners, and students how to read a paper critically rather than take its abstract on trust. That teaching heritage still shapes the tools. The questions are written in plain language, each comes with hint prompts, and the expected answers are “yes”, “no”, or “can’t tell”. This makes CASP the natural first instrument for a student-led review, a journal club, or a team appraising a mixed bag of designs before committing to heavier tools. It also makes CASP one of the most cited names in the wider family of critical appraisal tools mapped by study design, where it sits alongside more formal instruments rather than replacing them.

The trade-off for that accessibility is granularity. CASP tells you whether a study is broadly trustworthy and usable; it does not decompose bias into domains the way Cochrane instruments do. Knowing which side of that line your review sits on is the single most important decision before you appraise anything. It helps to remember what the checklists were designed to do. They emerged from workshops where a mixed room of clinicians, managers, and researchers had thirty minutes to decide whether a paper should change practice, and every design choice serves that setting: short question stems, hint prompts beneath each question suggesting what to look for in the paper, and an explicit instruction that the first questions are screening questions that can end the appraisal early. A tool built for a workshop transfers naturally to a review team, but only if the team adds the record-keeping a workshop never needed: written justifications, duplicate appraisal, and a documented reconciliation process.

The main CASP checklists

The programme maintains a separate checklist for each major study design, because the questions that expose weakness in a trial are not the ones that expose weakness in an interview study. The core set covers:

  • Randomised controlled trials: randomisation, allocation concealment, blinding, completeness of follow-up, and whether groups were treated equally apart from the intervention.
  • Systematic reviews: a clearly focused question, an adequate search, appraisal of the included studies, and whether pooling was reasonable.
  • Qualitative studies: appropriateness of the methodology, the recruitment strategy, the researcher-participant relationship, and the rigour of the analysis.
  • Cohort studies: recruitment of an unbiased cohort, measurement of exposure and outcome, and handling of confounding and follow-up.
  • Case-control studies: selection of cases and controls, exposure measurement, and confounders.
  • Diagnostic studies: an appropriate reference standard, blinded comparison, and a sensible patient spectrum.
  • Economic evaluations: a defined economic question, costed alternatives, and sensitivity of the conclusions to the assumptions.
  • Clinical prediction rules: derivation and validation of the rule in appropriate populations.

Recent revisions have refreshed the wording and added guidance notes, but the family logic is stable: pick the checklist that matches the design in front of you, never force one checklist across every design in the review. Version control matters more than it looks: the randomised trial checklist, for example, was revised in 2024 with reworded questions, so an appraisal citing “the CASP RCT checklist” without a year is not fully reproducible. Download the current forms from the programme directly, state the year in your methods, and archive the exact document your team used, because a reader attempting to audit your judgements needs the same question wording you saw.

The three-section logic: validity, results, applicability

Every CASP checklist follows the same arc. Section A screens for validity, beginning with two or three screening questions so fundamental that a “no” means it may not be worth continuing: was there a clearly focused question, and was the method appropriate to answer it? Section B then asks what the results are, how precise they are, and whether the analysis was sufficiently rigorous. Section C asks whether the results will help locally: can they be applied to your population, were all important outcomes considered, and do the benefits justify the harms and costs. This validity-results-applicability structure is why the checklists transfer so well between designs, and why they double as teaching instruments: they mirror the order in which an experienced reader actually interrogates a paper.

The screening questions deserve respect rather than a reflexive “yes”. If a study has no clearly focused question, the remaining answers become unanchored, because validity can only be judged relative to what the study set out to do. In a review context, a study that fails screening is rarely discarded on that ground alone, since it has already passed eligibility, but the failure should be recorded and should colour how much weight the study carries in the synthesis. Section C is the part review teams most often rush, and wrongly: for a reader deciding whether evidence transfers to their population, the applicability questions are the entire point of the exercise.

Why there is no numeric score, and why summing answers is wrong

CASP deliberately publishes no scoring system, and the temptation to invent one is the most common misuse of the tools. Counting the “yes” answers and reporting “8 out of 10” treats every question as equally important, which they are not: a trial that fails on allocation concealment is compromised in a way that a trial with a vague funding statement is not, yet a summed score weights them identically. A total also hides which weakness is present, so two studies with the same score can be flawed in completely different ways. The defensible approach is to report the answer to each question, add a short justification quoting the source text, and reach a reasoned overall judgement in prose. If a journal or supervisor insists on a summary, describe studies as raising no, minor, or major methodological concerns and say which questions drove that call, a distinction explored further in the difference between risk of bias and quality assessment.

Running a CASP appraisal step by step

In a review, the checklist works best as a fixed procedure applied identically to every study:

  1. Match each included study to its design-specific checklist at the point of data extraction, and record the checklist and version in the methods section.
  2. Have two appraisers pilot the checklist on two or three studies first, agreeing on how strictly to interpret each question before scaling up.
  3. Answer the screening questions first, then every remaining question as “yes”, “no”, or “can’t tell”, quoting the page or table that supports each answer.
  4. Write a short narrative judgement per study, naming the questions that raised concern and stating how much the concerns matter for the review’s question.
  5. Reconcile the two appraisers’ answers, resolving disagreement by discussion or a third opinion, and keep the pre-reconciliation answers so the process is auditable.

None of this is unique to CASP, but the checklists’ informality tempts teams to skip it. The tool is only as reproducible as the procedure wrapped around it.

CASP versus formal risk-of-bias tools

CASP is a critical appraisal framework, not a formal risk of bias instrument, and the distinction has practical teeth. Cochrane reviews and an increasing number of journals expect randomised trials to be judged with the RoB 2 signalling-question algorithm and non-randomised intervention studies with ROBINS-I and its target-trial logic, because those tools produce domain-level judgements that feed directly into GRADE. A CASP appraisal cannot be converted into those judgements after the fact. So the working rule is: for an intervention review destined for a methods-conscious journal, use the Cochrane instruments for bias and keep CASP for teaching, scoping, or first-pass triage; for reviews of qualitative research, or reviews where no domain-based tool exists for the design, CASP is a respected, citable choice. Observational designs sit in between, with many teams preferring the Newcastle-Ottawa Scale for cohort and case-control studies or a design-matched checklist from the Joanna Briggs Institute suite when they want something closer to a structured appraisal per design.

CASP for qualitative evidence synthesis

The qualitative checklist is where CASP is genuinely dominant. It is the most widely used appraisal instrument in qualitative evidence synthesis, cited across meta-ethnographies, thematic syntheses, and framework syntheses, and it is the tool Cochrane’s qualitative methods guidance names most often. Its ten questions probe the things that matter for trustworthiness in qualitative work: whether the design fits the question, whether recruitment and data collection were justified, whether the researcher’s own role was examined (reflexivity), whether ethics were addressed, and whether the analysis was sufficiently rigorous and grounded in the data. Appraisal here rarely excludes studies outright; more often it informs how much each study contributes to the synthesis and supports approaches such as GRADE-CERQual when rating confidence in review findings. Because judgements about analytical rigour depend on understanding how the primary authors worked with their data, it pairs naturally with how coding works in qualitative reviews. Teams running a full synthesis under deadline often hand this stage to our qualitative evidence synthesis service, where appraisal and synthesis are done by the same methodologists.

Reporting CASP results in a review

Report the appraisal the same way you would report screening: transparently and per study. Good practice is a table with one row per study and one column per CASP question, cells showing “yes”, “no”, or “can’t tell”, plus a comments column citing the passage that supports each judgement. State in the methods section which checklist version was used, that two reviewers appraised independently, and how disagreements were resolved; agreement can be quantified the same way as inter-rater reliability at the screening stage. Then make the appraisal do work: discuss whether the weaker studies change the synthesis, rather than filing the table as an appendix ornament. A review that appraises with CASP and never mentions the results again has performed a ritual, not an appraisal. Handled properly, the table becomes part of the audit trail that reporting standards such as PRISMA 2020 expect, and it gives readers the evidence behind every judgement you made.

The summary is simple. Match the checklist to the design, answer every question with a citation, resist the urge to add numbers that the tool’s own authors never intended, and switch to a formal domain-based instrument when the review’s venue demands one. Used that way, the CASP family remains what it was built to be: the clearest entry point into critical appraisal that exists.