An individual patient data meta-analysis, usually shortened to IPD meta-analysis, is a synthesis in which the reviewers obtain the raw, participant-level data from each eligible study and re-analyse everything centrally, instead of pooling the summary results printed in the papers. Each trial’s original records, one row per patient with treatment allocation, baseline characteristics, and outcomes, are collected under data-sharing agreements, cleaned, harmonised, and analysed to a single standardised analysis plan. It is widely regarded as the gold standard of evidence synthesis, and it earns that label by answering questions that aggregate results simply cannot reach.

What raw data buys you that published summaries cannot

The conventional approach, pooling published effect estimates through the standard aggregate data meta-analysis workflow, inherits every choice the original trialists made: their outcome definitions, their follow-up cut-offs, their handling of missing data, their adjustment sets. Individual patient data removes that inheritance, and four advantages follow.

First, standardised analysis. The central team can define outcomes identically across trials, apply the same intention-to-treat rules, impute missing data consistently, and use one statistical model everywhere, so differences between studies reflect the evidence rather than analytic fashion. Second, patient-level effect modifiers. Aggregate data can only relate treatment effects to trial-level averages, an approach vulnerable to ecological bias and weak even when done carefully through subgroup analysis and meta-regression; with participant-level data you can test directly whether the effect differs by age, sex, baseline severity, or biomarker status, using within-trial information with far greater power. Third, time-to-event precision. Rather than pooling reported hazard ratios from survival analyses computed at whatever follow-up each paper happened to report, the team re-runs the survival analysis from raw event times, often with longer, updated follow-up supplied by the trialists. Fourth, checking randomisation integrity. With the raw data in hand, reviewers can verify baseline balance, hunt for impossible dates or duplicated patients, and confirm that the analysis matches the allocation, checks that have exposed flawed and even fraudulent trials that aggregate reviews had accepted at face value.

The collaborations that defined the method

The approach was not invented on a whiteboard; it was forged by a handful of long-running trialist collaborations whose results reshaped clinical practice. The Early Breast Cancer Trialists’ Collaborative Group, running since the 1980s, has periodically re-analysed participant-level data from hundreds of trials and its findings on tamoxifen and chemotherapy changed breast cancer treatment worldwide. Similar standing collaborations in hypertension, cholesterol lowering, and perinatal medicine demonstrated the same pattern: pooled raw data, refreshed follow-up, and patient-level subgroup answers that no aggregate synthesis could have produced. These groups also established the social model the method still relies on, a secretariat holding the data under strict agreements, trialists as named collaborators, and analyses agreed collectively before anyone sees results, which is why the label collaborative meta-analysis is often used interchangeably with the individual patient data label. The model has since spread from oncology and cardiovascular medicine into mental health, obstetrics, and public health, and the number of published individual participant data projects has grown steadily year on year, helped by journals and funders that increasingly expect trial data to be shareable as a condition of publication or funding.

One-stage versus two-stage modelling

Once the data are assembled there are two broad modelling routes. The two-stage approach analyses each trial separately using the standardised plan, producing one effect estimate per study, and then combines those estimates exactly as a conventional synthesis would, choosing between fixed-effect and random-effects pooling in the second stage. It is transparent, familiar, and preserves randomisation within each trial by construction. The one-stage approach fits a single hierarchical model to all patients from all trials simultaneously, with trial-specific terms to respect clustering. The two usually agree for overall effects, and where they diverge the cause is typically a hidden modelling difference rather than magic in either method. The one-stage model earns its keep with sparse data, rare events, and above all treatment-covariate interactions, where it can separate within-trial from across-trial information cleanly. Whichever route is chosen, the decision belongs in the statistical analysis plan before the data arrive, and any exploration of between-study heterogeneity should be specified there too.

Two practical cautions apply to interaction analyses in particular. Effect modifiers, like every other analysis choice, must be pre-specified, because with raw data the temptation to trawl dozens of patient characteristics until something turns significant is stronger than ever, and post hoc subgroup findings from participant-level data are no more trustworthy than post hoc findings anywhere else. And continuous modifiers such as age or baseline risk should be modelled continuously rather than split at arbitrary cut-points, which discards information and can manufacture spurious subgroup effects. Handled with that discipline, the interaction analysis is the single strongest argument for collecting raw data at all, because it is the analysis aggregate results are structurally unable to deliver.

The data collection reality: agreements, collaboration, and years

The statistics are the easy half. The defining labour of an individual patient data project is getting the data. Each trial group must be identified, contacted, and persuaded; legal teams on both sides negotiate data-sharing agreements covering permitted uses, security, and publication rights; ethics and privacy rules differ by country; and the data, once received, arrive in a dozen incompatible formats with undocumented variable names that must be decoded with the original trialists’ help. This is why serious projects are structured as collaborative groups, with trial investigators credited as co-authors and consulted on design, an arrangement that both raises supply rates and improves the analysis plan. Repositories and sponsor platforms for clinical trial data have eased access for newer industry trials, but older and academic trials still depend on personal persuasion of the kind described in our guide to contacting authors for missing data, sustained over a much longer campaign. Realistic timelines run two to five years from protocol to publication, and funding applications should say so plainly.

From protocol to publication, step by step

In outline, a well-run project moves through seven stages. First, a published protocol and statistical analysis plan fix the questions, the eligibility criteria, and the models before any data arrive. Second, systematic searching identifies every eligible trial, published or not, since unpublished trials can and should contribute data even without a paper. Third, the invitation phase recruits trial groups into the collaboration and negotiates the agreements. Fourth, incoming datasets pass through checking and harmonisation: variables are recoded to a common dictionary, ranges and dates are validated, randomisation and follow-up are verified, and queries go back to the trialists until each dataset is signed off. Fifth, the frozen master dataset is analysed exactly as the plan specified. Sixth, results are circulated to the collaboration for challenge before anything is submitted. Seventh, the manuscript reports to the PRISMA-IPD standard with the full data flow on display. The sequence looks like a conventional review with extra steps, but the centre of gravity is different: most of the calendar time and most of the budget sit in stages three and four, not in the analysis.

Non-supplied datasets and availability bias

No project obtains everything. Trials are lost because investigators have retired, sponsors refuse, data were destroyed, or agreements stall, and retrieval rates in published projects commonly land between sixty and ninety percent of eligible participants. The danger is data availability bias: if the trials that supply data differ systematically from those that do not, for instance if unfavourable trials are quietly withheld, the synthesis inherits a distortion no amount of modelling removes. The defences are procedural and analytic. Procedurally, teams document every request and refusal so the flow of trials is fully auditable. Analytically, they compare the characteristics and published results of supplied and non-supplied trials, and run sensitivity analyses combining the individual-level data with published aggregate results from the missing trials to test whether conclusions survive. A project that reports an effect from seventy percent of the evidence without examining the other thirty percent has not finished its analysis.

Reporting with PRISMA-IPD

Individual patient data projects report to PRISMA-IPD, a dedicated extension of the PRISMA reporting guideline published for exactly this design. Beyond the familiar items, it requires reporting of how data were sought and obtained, trial by trial; how the supplied datasets were checked and harmonised, including verification of randomisation and follow-up; which trials could not contribute and how that risk was assessed; and how the one-stage or two-stage models were specified. The flow diagram tracks participants as well as studies, so readers can see what fraction of the theoretically available evidence the analysis actually contains. Meeting the checklist is not bureaucracy; it is the audit trail that lets a reader weigh the project’s central claim of being the gold standard. Reviewers and editors read these sections closely because the label attracts prestige, and a project that obtained data from a third of the eligible trials while analysing them beautifully is not the gold standard of anything; the checklist is what forces that arithmetic into the open.

When aggregate meta-analysis remains the sensible choice

The gold standard is not the universal standard. An aggregate synthesis remains the right call when the review question concerns the overall effect and the published results are consistent, well reported, and based on comparable outcomes, because collecting raw data would add years and cost for the same answer. It is also the pragmatic choice when decision makers need an answer soon, when the evidence base is young and still moving, or when data custodians realistically will not share. The considered strategy many groups follow is sequential: run a rigorous aggregate review first, and escalate to individual patient data only if it surfaces the specific problems raw data solve, such as credible effect modification, inconsistent outcome definitions, or doubts about the integrity of influential trials. Used that way, the method is reserved for the questions that justify its cost, where who benefits matters as much as whether anyone does, and there it has no substitute.

There is also a middle path worth knowing about. Hybrid designs combine participant-level data from the trials that supply it with published aggregate results from those that do not, keeping the whole evidence base in view while exploiting the raw data where it exists. And for questions about prognosis or diagnostic accuracy rather than treatment effects, the same machinery of collection, harmonisation, and hierarchical modelling applies, which is why individual participant data methods now extend well beyond randomised trials. The decision, in the end, is an economic one about information: pay the price of raw data when the question demands answers only raw data can give, and spend the savings on doing the aggregate synthesis properly when it does not.