GRADEpro GDT (Guideline Development Tool) is the official web application for applying the GRADE approach: it takes the pooled results of a systematic review, walks you through the structured judgements about how trustworthy each result is, and produces the Summary of Findings table that presents effects and certainty of evidence together. Built by Evidence Prime with the GRADE Working Group and used for Cochrane reviews, health technology assessments, and clinical guidelines, it is the standard route from a forest plot to a defensible certainty rating.
Where GRADEpro sits in the review workflow
By the time you open GRADEpro, the analysis should be finished. The software does not pool anything; it consumes the outputs of your synthesis, the pooled estimate, the confidence interval, the number of studies and participants, and turns them into a rated evidence profile. Its place in the sequence is therefore after data extraction, after appraisal, and after the meta-analysis itself, whether that was run in RevMan Web or a statistical package. The intellectual framework it implements is the GRADE system of rating evidence as high, moderate, low, or very low certainty, which we unpack in our guide to GRADE and the certainty of evidence. GRADEpro’s contribution is discipline: it makes you record a judgement and a reason for every domain of every outcome, which is exactly what peer reviewers check. It is worth being clear about what the ratings mean, because the language trips people: certainty is a property of the body of evidence for one outcome, not of the review, so a single review routinely reports high certainty for mortality and low certainty for quality of life side by side, and that spread of judgements is the finding, not a defect.
Building a Summary of Findings table step by step
A workable sequence looks like this. First, create a project and define the question in population, intervention, comparison format; the software supports intervention questions as well as diagnostic and prognostic ones. Second, add up to around seven critical and important outcomes, the ones decision makers actually need, not every outcome the review extracted. Third, for each outcome enter the study numbers, participant totals, and pooled effect. Fourth, work across the certainty domains, recording a judgement and an explanatory footnote for each. Fifth, choose the table format and export to Word or into the review. The selection of outcomes deserves real thought before the software is even open, because a table crowded with trivial endpoints buries the result a clinician came to find. It also pays to fix the outcome list before results are known: choosing which outcomes to display after seeing which ones favour the intervention is a form of selective reporting, and a protocol that names the critical outcomes in advance is the defence.
Entering pooled estimates and absolute effects
For a dichotomous outcome you enter the relative effect, typically a risk ratio or odds ratio with its confidence interval, and GRADEpro converts it into absolute effects: the number of events per 1,000 people in the comparison group and the corresponding number with the intervention. You supply the baseline risk, usually from the control arms or from an external source when control-arm risk is unrepresentative, and the software does the arithmetic and phrases the difference. This absolute translation is the entire point of the table: a risk ratio of 0.80 means something quite different at a baseline risk of 2% than at 40%. If the distinction between the relative measures is hazy, our piece on odds ratios versus risk ratios settles it, and the related idea of translating effects into patient numbers is covered in our guide to the number needed to treat. Continuous outcomes are entered as mean differences or standardised mean differences with an anchor that tells readers what the scale means, ideally alongside a minimal important difference so the reader can judge whether a statistically clear change is clinically worth having. Whatever the outcome type, enter the numbers exactly as the synthesis produced them and let the software do the conversions; rounding by hand at this stage is how tables end up contradicting their own forest plots.
Evidence profiles versus Summary of Findings tables
GRADEpro produces two related outputs, and knowing which your document needs saves rework. The evidence profile is the long format: it shows every certainty judgement explicitly, one column per domain, so a reader can see exactly why an outcome sits at moderate rather than high certainty. The Summary of Findings table is the compact format: it presents the effects, the participant numbers, the overall certainty grade, and the footnotes, without the domain-by-domain columns. Guideline panels and methods-heavy journals usually want the profile; systematic reviews usually present the Summary of Findings table and keep the profile as supplementary material. GRADEpro generates both from the same underlying judgements, so the sensible workflow is to complete the assessment once and export whichever formats the venue requires, rather than maintaining two versions by hand.
The eight GRADE domains inside the software
For each outcome, GRADEpro presents five reasons to downgrade certainty and three to upgrade it. The five downgrades are risk of bias (limitations in the included studies), inconsistency (unexplained variation between study results, judged partly through the statistics discussed in our guide to heterogeneity in meta-analysis), indirectness (evidence from populations, interventions, or outcomes that differ from the question), imprecision (confidence intervals too wide to support a single course of action), and publication bias (suspicion that missing studies would change the answer, explored in our article on publication bias in systematic reviews). Each is rated not serious, serious, or very serious. The three upgrades, applied mainly to evidence from non-randomised studies, are a large effect, a dose-response gradient, and plausible confounding that would work against the observed effect. Ratings start at high for randomised trials and low for observational evidence, then move with the judgements.
Making the judgements well
The software records judgements; it does not make them, and the quality of a table rests on how the domains are applied. For risk of bias, the question is not whether any included study had flaws but whether the studies carrying most of the weight in the pooled analysis did, which means reading the appraisal alongside the forest plot rather than counting red circles. For imprecision, current GRADE guidance asks whether the confidence interval crosses a threshold that would change a decision, not merely whether it crosses the line of no effect. For inconsistency, an explained pattern, for example a difference by dose that a pre-specified subgroup analysis accounts for, is not a reason to downgrade, while unexplained scatter is. And indirectness is judged against the review question, so evidence from a narrower population than the one you asked about can be serious even when every study is internally sound. These are exactly the judgement calls that separate an assessment which survives peer review from one that gets sent back.
Footnoting judgements properly
The difference between a credible table and a decorative one is the footnotes. Every downgrade, and every decision not to downgrade where a reader might expect one, should carry an explanation: which trials were at high risk of bias and in which domains, what the I-squared was and why the inconsistency could not be explained, how far the confidence interval crosses the decision threshold. GRADEpro makes this mechanically easy, attaching lettered footnotes to any cell, but it cannot write the reasoning for you. Write footnotes as you make each judgement, not retrospectively, and keep them specific enough that a methodologist could reconstruct the decision from the note alone.
Importing from RevMan and working with Cochrane reviews
For Cochrane authors, GRADEpro connects directly to RevMan so that pooled results flow across without retyping, and the finished Summary of Findings table travels back into the review. The link matters for accuracy as much as speed: transcription errors between an analysis and its summary table are one of the commonest problems editors catch in a Cochrane review. Outside that ecosystem you can still work smoothly by entering pooled results manually from any analysis output, and the export formats, Word tables and interactive versions, drop into most journal submissions without rework.
Evidence-to-decision frameworks
GRADEpro is more than a table builder. Its evidence-to-decision framework module supports the step from evidence to recommendation, prompting a guideline panel to weigh the balance of benefits and harms, the certainty ratings, patient values and preferences, resource use, equity, acceptability, and feasibility, and to record how each consideration pushed the recommendation. Systematic reviewers may never need this module, but for guideline groups it is the part that turns rated evidence into a transparent recommendation with an audit trail, and it is why the tool carries Guideline Development Tool in its name. The same module underpins collaborative features in the paid tiers, where panel members vote remotely on judgements and the software collates the responses, an arrangement that became standard for international guideline groups that can no longer convene everyone in one room.
Common mistakes to avoid
The recurring failures are predictable. Unexplained downgrades: a rating drops a level with no footnote, leaving readers to guess the reason. Missing absolute effects: the table shows only relative effects, which misleads whenever baseline risk is low. Double counting, where the same wide confidence interval is punished under both inconsistency and imprecision without justification. Rating the study rather than the outcome, when GRADE judgements are outcome-specific by design. Mechanical rules applied without judgement, such as downgrading for any I-squared above 50% regardless of whether the inconsistency affects the conclusion. And tables that include every extracted outcome instead of the handful that matter. Each mistake is cheap to avoid at the time and expensive to fix after peer review. A final habit worth building: draft the GRADE assessment while the synthesis is fresh, ideally the same week the pooled analyses are finalised, because judgements about inconsistency and imprecision are far easier to write when the reasons are still in working memory than when a reviewer report arrives months later.
Free versus paid access
At the time of writing, GRADEpro operates a freemium model. The standard browser version is free for individual researchers and small projects, and that free tier is how most systematic reviewers use it, including for building Summary of Findings tables. Paid team and enterprise licences add collaboration features, panel voting, project management, and administrative controls aimed at guideline programmes and organisations, with team pricing at a level pitched at funded projects rather than students. Tiers and allowances change, so confirm the current arrangement on the GRADEpro pricing page before budgeting. For a typical review team the practical answer is simple: the free version does everything the Summary of Findings table requires, and if you would rather hand the whole judgement process to specialists, our GRADE assessment service delivers the completed profile with the reasoning written out.