The Mixed Methods Appraisal Tool, usually cited as MMAT, is a single instrument for appraising the methodological quality of the five study categories a mixed studies review can contain: qualitative research, quantitative randomised controlled trials, quantitative non-randomised studies, quantitative descriptive studies, and mixed methods studies. The current 2018 version opens with two screening questions, then applies five methodological quality criteria matched to each study’s design, and deliberately provides no overall numeric score.

Why mixed studies reviews needed one tool

Reviews that combine qualitative and quantitative evidence face an appraisal problem no single-design review has: three or four different instruments, with incompatible response formats and outputs, running side by side. A team might apply a trial tool, an observational checklist, and a qualitative checklist in one review and then struggle to say anything coherent about quality across the whole evidence base, because the outputs of those instruments share neither a scale nor a vocabulary. The MMAT, first piloted in 2009 by Pierre Pluye and colleagues at McGill University and substantially revised in 2018 through a literature review and an expert e-Delphi study, was built to close that gap: one tool, one response format, one appraisal vocabulary across every included design. It has become the default appraisal instrument for mixed studies reviews in health and social research, and it occupies a distinctive niche among the major critical appraisal tools organised by design: not the deepest instrument for any single design, but the only one that spans all five categories at once.

The 2018 revision is worth dating precisely because pre-2018 papers describe a different tool. The 2011 version carried a percentage quality score and fewer criteria per design; the 2018 version was rebuilt after a review of critical appraisal instruments and a two-round e-Delphi with international mixed methods experts, settling on five criteria per category and dropping the score. The revision also produced a criteria manual with indicators and worked examples for every rating, which is the document appraisers should work from rather than the one-page grid alone. Reviews citing “the MMAT” should therefore name the version and cite the 2018 user guide, and appraisers moving from an older protocol must not mix criteria sets across studies.

The two screening questions

Before any design-specific criteria, the MMAT asks two gatekeeping questions of every paper: are there clear research questions, and do the collected data allow those questions to be addressed? These exist to filter out papers that are not empirical studies at all, editorials, theoretical pieces, protocols, methods papers, because a study that fails screening cannot meaningfully be appraised further. A “no” or “can’t tell” on either question is a signal to reconsider whether the record should have survived full-text eligibility screening in the first place, which is why many teams fold the two questions into their eligibility form rather than discovering the problem at appraisal stage. Note what the screening questions do not do: they do not judge quality, only appraisability. A weak empirical study passes screening and then earns its weak ratings on the criteria; a brilliant conceptual paper fails screening not because it is bad but because the MMAT has no yardstick for it. Keeping that distinction clear prevents the screening stage from becoming an informal quality filter that no protocol ever described.

Five criteria for each of five designs

The 2018 revision gives every design category exactly five criteria, answered yes, no, or can’t tell. The symmetry is deliberate: five criteria per category keeps the appraisal burden equal across designs, so no study type is judged more harshly simply because its checklist happened to be longer. The criteria themselves were selected in the revision process for being the strongest available markers of methodological quality in each tradition, which is why they read as a distillation of the larger design-specific tools rather than a new invention:

  • Qualitative studies: appropriateness of the qualitative approach, adequacy of data collection, whether findings are derived from the data, whether the interpretation is grounded in the data, and coherence across question, methods, and analysis.
  • Quantitative randomised controlled trials: appropriate randomisation, comparable groups at baseline, completeness of outcome data, blinding of outcome assessors, and adherence to the assigned intervention.
  • Quantitative non-randomised studies: representativeness of participants, appropriate measurements of outcome and exposure, complete outcome data, accounting for confounders, and intervention delivery as intended.
  • Quantitative descriptive studies: relevance of the sampling strategy, representativeness of the sample, appropriate measurements, acceptable nonresponse bias, and appropriate statistical analysis.
  • Mixed methods studies: an adequate rationale for mixing, effective integration of the qualitative and quantitative components, adequate interpretation of the combined outputs, attention to divergences between components, and adherence of each component to the quality criteria of its own tradition.

The mixed methods category is the demanding one: a mixed methods study is appraised on its five integration criteria and on the criteria of each component design, fifteen criteria in all, on the premise that a study can be no stronger than its weakest component.

Selecting the right category is itself a judgement call that should be made before appraisal begins. A study that collects open-ended survey comments alongside numeric items is not automatically a mixed methods study; the label requires a genuine qualitative component with its own question, data collection, and analysis, integrated with the quantitative strand. Many self-described mixed methods papers are better appraised as quantitative descriptive studies with anecdotal quotation, and saying so, with reasons, is part of an honest appraisal. The same care applies at the quantitative boundary: a before-and-after evaluation without a comparison group belongs in the non-randomised category, not the trial category, and a secondary analysis of routine data is descriptive unless it formally tests an association. Classification drives which five criteria a study faces, so a wrong category quietly invalidates the appraisal that follows it.

Why the authors discourage an overall score

Earlier MMAT versions allowed a percentage quality score, and the 2018 revision pointedly removed it. The reasoning matches the other modern appraisal instruments: a summed score treats all criteria as equally important, hides which weakness a study actually has, and lets strong reporting in trivial areas launder a serious flaw. The authors instead ask reviewers to report the rating on each criterion and to describe quality in words, exactly the logic that separates judgement-based tools from score-based ones in the risk of bias versus quality assessment debate. If a summary is unavoidable, the sanctioned move is descriptive, reporting how many criteria each study met while stating that criteria are not equivalent, never a percentage presented as a quality grade, and never an exclusion threshold invented after the appraisal. Peer reviewers have caught up with this point: manuscripts reporting “MMAT scores” as percentages now routinely attract revision requests citing the tool’s own user guide, so the shortcut costs more time than the per-criterion table it was meant to avoid.

Running the MMAT step by step

Applied inside a review, the tool rewards the same procedural discipline as any other appraisal instrument:

  1. Classify every included study into one of the five categories, in duplicate, using pre-agreed definitions, and record the classification with its rationale.
  2. Apply the two screening questions; a failure here sends the record back to the eligibility discussion rather than onwards to appraisal.
  3. Rate the five design-matched criteria as “yes”, “no”, or “can’t tell”, each with a citation to the supporting passage, remembering that mixed methods studies also take the criteria of their component designs.
  4. Reconcile the two appraisers’ ratings, keeping the original answer sets, and record overall comments per study in prose rather than a number.
  5. Decide, against the protocol, what the ratings change: weighting in the synthesis, sensitivity analyses, or the framing of confidence in the findings.

Using the MMAT inside a mixed studies review

Appraisal only matters if it changes something downstream. In a convergent synthesis, MMAT ratings inform how much interpretive weight each study’s findings carry; a theme supported only by studies that failed their qualitative criteria deserves flagging, which requires the appraiser to understand how the primary analyses were built, the territory covered in coding data in qualitative reviews. In sequential designs, where a quantitative synthesis and a qualitative synthesis run separately before integration, MMAT results can drive a sensitivity analysis on the pooled quantitative estimate and a parallel confidence assessment on the qualitative side. The method section should state that two reviewers appraised independently, how conflicts were reconciled, and what the appraisal changed; agreement between appraisers is worth quantifying, since a low rate signals the criteria were being interpreted differently and the calibration step needs repeating. Teams synthesising substantial qualitative components often pair the MMAT stage with our qualitative evidence synthesis service, where the same methodologists carry appraisal through to the integrated findings.

MMAT versus separate design-specific tools

The honest comparison is depth against coherence. Design-specific instruments go deeper: the five trial criteria in the MMAT are a compressed echo of the domains that RoB 2 examines per outcome, and its five qualitative criteria are leaner than the ten questions in the CASP qualitative checklist. A review whose conclusions hang on a meta-analysis of trials should use the Cochrane tool for those trials, full stop. But a mixed studies review appraising fifteen studies across four designs with four different instruments will produce four incommensurable sets of judgements and no way to discuss quality across the review as a whole. The pragmatic pattern used by experienced teams: apply the MMAT review-wide for a common quality vocabulary, and add a deeper design-specific tool for the component that carries the heaviest inferential load. Either way, declare the choice in the protocol, appraise in duplicate, publish the per-criterion table, and let the appraisal visibly shape the synthesis. That, not the choice of instrument, is what separates a credible mixed studies review from a decorative one.

A final word on reporting, because this is where otherwise sound appraisals are undone at peer review. The MMAT’s own guidance asks for three things in the manuscript: a methods statement naming the 2018 version, the number of independent appraisers, and the reconciliation process; a table presenting the rating on every criterion for every study; and a narrative that describes the quality of each study category in words. Reviews that instead report a lone sentence, “quality was assessed with the MMAT and most studies were of moderate quality”, discard the information the tool generated and invite exactly the reviewer criticism the appraisal was meant to prevent. The per-criterion table costs half a page in a supplement and answers every question a sceptical reader might ask about what “moderate” means. In appraisal, as in the rest of a review, transparency is cheaper than authority and far more convincing.