A levels of evidence hierarchy ranks study designs by how much protection each offers against bias, running from systematic reviews of randomised trials at the top to expert opinion at the bottom. Nursing uses several competing hierarchies, they do not agree with each other, and the number you assign to a study is meaningless unless you name which one you used.
There is no single hierarchy, and the differences are real
Candidates often write that a study is Level III as though that were a property of the study. It is not. It is a property of the study under a named scheme.
The scheme bundled with the Johns Hopkins framework compresses designs into three broad levels and then rates quality separately. Melnyk and Fineout-Overholt use seven levels. JBI publishes design-specific hierarchies that differ by question type, so effectiveness, prevalence, and meaning questions each get their own ladder. A cohort study can legitimately sit at different numbers in each.
This is why the number alone communicates nothing. Reporting practice should always be the scheme, its edition, and then the level.
Level is not quality, and conflating them is the core error
A hierarchy ranks what a design is capable of, assuming it was executed properly. It says nothing about whether this particular study was executed properly.
A randomised trial with concealed allocation failures and forty percent attrition is still a randomised trial by design, and still sits high on a design hierarchy. A well-conducted cohort study with complete follow-up may be far more trustworthy. Treating the level as a proxy for trustworthiness inverts the actual evidence.
This is exactly why the Johns Hopkins scheme separates level from quality, and why risk of bias assessment is a distinct exercise from assigning a level. The hierarchy tells you where to start; the appraisal tells you what to conclude.
Where hierarchies break down
Design hierarchies were built around questions of intervention effectiveness, and they degrade badly outside that.
For a question about lived experience, a randomised trial is not a superior form of evidence, it is the wrong form. Applying an effectiveness ladder to qualitative work produces the absurd result that every included study is low level. JBI addresses this by publishing separate hierarchies per question type, which is the coherent response. For synthesis of qualitative evidence, see meta-synthesis.
Diagnostic accuracy, prevalence, and prognosis questions each have their own appropriate designs, and a cross-sectional study is the right design for prevalence rather than a weak substitute for a trial.
How this relates to GRADE
GRADE is frequently described as a levels-of-evidence system and it is not one. A hierarchy classifies individual studies by design. GRADE assesses a body of evidence for a specific outcome, starting from design and then moving the rating up or down for risk of bias, inconsistency, indirectness, imprecision, and publication bias.
So a body of randomised trial evidence can end at low certainty under GRADE, and a body of observational evidence can be upgraded. Journals publishing intervention reviews increasingly expect GRADE rather than a design hierarchy, while local practice projects usually expect the hierarchy their framework bundles. Knowing which your audience wants avoids doing both badly.
Presenting levels in an evidence table
An evidence table carrying levels should state the scheme and edition in the caption or a footnote, give the level and the appraisal outcome in separate columns, and never sum levels into a score. Summing implies the intervals between levels are equal and comparable, and they are neither.
Where the table is feeding a formal synthesis rather than a local project, the appraisal column should point to the instrument used, drawn from the appraisal tools available for each design.