Risk of bias and quality assessment services

We apply the right appraisal tool to every study, assess in duplicate with documented judgements, and carry the result through to a defensible GRADE certainty rating.

4.9 / 5 across 1,194+ delivered projects

  • Cochrane and PRISMA 2020 methods
  • PhD methodologists
  • 100% human-written, no generative AI
  • Reproducible R and Stata code
  • Free quote within 24 hours
  • Mutual NDA on request

Risk of bias assessment is the structured appraisal of how much each included study’s design and conduct could distort its results. Our service applies RoB 2 to randomised trials, ROBINS-I to non-randomised studies, the Newcastle-Ottawa scale to observational designs, and AMSTAR 2 to existing reviews, with two assessors working independently, and feeds the result into a GRADE certainty rating for each outcome.

Why bias assessment is what reviewers scrutinise most

Two reviews can include the same studies and reach opposite levels of confidence, and the difference is usually the quality appraisal. A pooled estimate built on high-risk studies says far less than the same estimate built on well-conducted trials, and an editor will check that you have made that distinction visible rather than burying it. Risk of bias is judged on whether the right tool was used, whether the judgements are supported by evidence from the papers, and whether the result actually changes how the findings are interpreted. We treat it as analysis, not box-ticking, anchored in the principles set out in our guide to risk of bias assessment.

Choosing the correct instrument is the first decision and the one most often got wrong. We map each study to its design and apply the RoB 2 tool for randomised trials, the ROBINS-I tool for non-randomised studies, and the Newcastle-Ottawa Scale for observational cohorts and case-control studies, so the appraisal fits what was actually done.

What the assessment service covers

  1. 1

    Tool selection by design

    We match each included study to RoB 2, ROBINS-I, Newcastle-Ottawa, or AMSTAR 2 based on its design, and justify the choice in the methods.

  2. 2

    Independent duplicate appraisal

    Two trained assessors rate each study independently across every domain, then reconcile, with a third assessor resolving any deadlock.

  3. 3

    Supported domain judgements

    Each domain judgement is backed by a quote or specific reason from the paper, so every rating can be traced and defended.

  4. 4

    GRADE certainty rating

    We summarise the bias picture per outcome and produce a GRADE certainty of evidence table that feeds directly into your conclusions.

The appraisal tools we are trained in

RoB 2 for randomised trials

Domain-level appraisal of the randomisation process, deviations, missing data, measurement, and selective reporting.

ROBINS-I for non-randomised studies

Appraisal that adds confounding and selection domains randomisation would otherwise handle, for cohort and intervention studies.

Newcastle-Ottawa and JBI

Scales for observational cohort, case-control, and prevalence designs, matched to what each study actually did.

AMSTAR 2 and GRADE

Critical appraisal of existing reviews and a certainty-of-evidence rating carried through to your conclusions.

What you receive

You receive a completed assessment for every study, a per-domain record with the supporting evidence, a reconciliation log, and a GRADE certainty of evidence summary for each outcome. Where your review appraises existing reviews, we apply the AMSTAR 2 critical appraisal checklist. The assessment is structured to plug straight into the statistical analysis and synthesis and into the full systematic review write-up, so your bias ratings and your pooled results tell one consistent story.

  • A completed appraisal for every included study with the correct tool
  • A per-domain record citing the supporting quote or reason for each judgement
  • A traffic-light or summary figure ready for your results section
  • A reconciliation log showing how assessor disagreements were settled
  • A GRADE certainty-of-evidence table for each outcome
  • A drafted risk-of-bias paragraph for your methods and results

Who needs an independent appraisal

Reviewers needing a second assessor

Authors who need an independent second rater so the appraisal meets the dual-assessment standard reviewers expect.

Mixed-design reviews

Reviews including both randomised and non-randomised studies that need more than one appraisal tool applied correctly.

Guideline and certainty work

Teams that need a defensible GRADE certainty rating to support clinical recommendations.

Authors answering reviewers

Researchers asked to add a formal appraisal or a GRADE table before a manuscript can be accepted.

How a bias assessment is quoted

We quote a fixed fee shaped by the number of included studies, the mix of study designs and therefore the number of tools required, and whether you need a full GRADE certainty assessment across your outcomes. Send your included studies and we will return a fixed quote for a duplicate, fully documented appraisal.

Appraising studies yourself versus trained assessors

ConsiderationDoing it yourselfWorking with us
Tool choiceUsing one tool for every design undermines the whole appraisalEach study matched to the instrument its design requires
Judgement supportDomain ratings asserted without traceable evidenceEvery judgement backed by a quote or specific reason from the paper
IndependenceA single rater's judgement is hard to defendTwo assessors rate independently and reconcile, with a tie-breaker
Certainty ratingGRADE applied loosely, weakening the conclusionsA bias picture carried cleanly into a defensible GRADE table

Frequently asked questions

Which risk of bias tool should I use?
The tool follows the study design. RoB 2 is the standard for randomised trials, ROBINS-I is used for non-randomised studies of interventions, and the Newcastle-Ottawa Scale is common for observational cohort and case-control studies. AMSTAR 2 appraises existing systematic reviews. We match each included study to the correct tool and explain the choice in your methods.
What is the difference between RoB 2 and ROBINS-I?
RoB 2 assesses bias in randomised controlled trials across domains such as the randomisation process and deviations from intended interventions. ROBINS-I assesses non-randomised studies and adds domains that randomisation would normally handle, most importantly confounding. They are not interchangeable, and using the wrong one undermines the appraisal.
How does risk of bias feed into GRADE?
GRADE rates the certainty of the evidence for each outcome, and risk of bias is one of the domains that can lower that certainty. If the studies behind an outcome carry serious limitations, GRADE downgrades the certainty rating. We assess each study, summarise the bias picture per outcome, and carry that judgement through into the GRADE certainty table.
Should two people assess risk of bias?
Yes. Risk of bias involves judgement, so two trained assessors should evaluate each study independently and reconcile their ratings, with a third resolving any deadlock. We assess in duplicate, record the supporting quote or reason for every domain judgement, and document how disagreements were settled.

Send us your included studies, and we will assess every one with the right tool and deliver a GRADE certainty table for a fixed quote.

Free quote within 24 hours. 100% human-written by PhD methodologists.

Methodology reviewed by

Dr Eleanor Whitfield, PhD

Lead Review Methodologist

Leads protocol design, screening, and reporting across the practice, with a focus on reviews that survive editorial scrutiny.

Why researchers bring in a PhD methodologist

80+

systematic reviews are published every day (Hoffmann et al., 2021)

67.3 weeks

average time to complete a review in-house (Borah et al., BMJ Open 2017)

~70%

of published reviews rate critically low on AMSTAR 2 quality appraisal

4.9 / 5

our client rating across 1,194+ delivered projects