Estimands in Meta-Analysis: When a Shared PICO Still Pools Different Questions
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Estimands in meta-analysis matter because two trials can share the same PICO and still estimate different treatment effects. The population matches. The drugs match. The endpoint and follow-up match. Yet one trial asks what happened after assignment, including discontinuation and rescue therapy, while another asks what would have happened had those events never occurred.
Pool those estimates without noticing the difference and the calculation may be flawless while the question is incoherent. PICO aligns the nouns. The estimand aligns the treatment question.
The Clean Example: Rescue Therapy Changes the Question
Imagine four randomized trials comparing two glucose-lowering treatments over 52 weeks. In every trial, some participants discontinue treatment or start rescue medication. Those events happen after assignment, affect the observed outcome, and complicate what “the treatment effect” means.
Treatment-policy question
What is the effect of assignment to each treatment strategy, regardless of later discontinuation or rescue? Those downstream events remain part of the practical treatment experience.
Hypothetical question
What would the effect have been if discontinuation or rescue had not occurred? This targets a counterfactual efficacy contrast and usually requires additional assumptions.
Neither question is universally superior. They serve different decisions. The error is to combine them as though they were noisy measurements of one automatic truth.
Interactive pooling audit
Same PICO, Different Treatment Questions
These four simulated trials share the same population, treatments, endpoint, and 52-week horizon. Negative values favor the intervention on a clinical-score scale. Change the target estimand to see why PICO alignment alone can produce a polished answer to a mixed question.
Trial A
Treatment policy-3.2 points
Follow-up continues after discontinuation or rescue therapy.
Included
Trial B
Hypothetical-7.1 points
Models the outcome had discontinuation or rescue therapy not occurred.
Included
Trial C
Treatment policy-4.0 points
Uses outcomes regardless of treatment discontinuation or rescue therapy.
Included
Trial D
Hypothetical-7.7 points
Targets efficacy in a world without discontinuation or rescue therapy.
Included
Illustrative fixed-effect summary
-5.5 points
95% interval -6.4 points to -4.7 points · 4 trials
The arithmetic is valid; the target is muddy. This summary averages pragmatic effectiveness with efficacy in a hypothetical world without discontinuation or rescue. Statistical compatibility does not make those questions interchangeable.
Teaching note: the values are simulated and the fixed-effect summaries are deliberately simple. Real synthesis also needs a justified heterogeneity model, compatible effect measures, uncertainty about classification, and careful review of whether trial-level analyses actually estimate the stated estimands.
What an Estimand Adds Beyond PICO
PICO remains indispensable for review scope. It describes who, which interventions, which comparison, and which outcome. The estimand framework adds the precision needed to understand the effect each trial is trying to estimate.
| Attribute | Extraction question | Pooling risk if vague |
|---|---|---|
| Population | Which patients does the effect describe? | Different eligibility or post-randomization subsets |
| Treatment conditions | What exactly is being compared? | Different versions, doses, or background care |
| Variable or endpoint | What outcome and time point define the contrast? | Nominally similar but clinically different outcomes |
| Intercurrent-event strategies | How are discontinuation, rescue, switching, or death reflected in the question? | Effectiveness and hypothetical efficacy mixed together |
| Population-level summary | Difference in means, risk ratio, hazard ratio, or another summary? | Different scales or summaries treated as equivalent |
Five Strategies Are Not Five Analysis Methods
ICH E9(R1) describes treatment policy, hypothetical, composite, while-on-treatment, and principal-stratum strategies for reflecting intercurrent events in the clinical question. These labels describe the target effect, not a menu of statistical commands. A mixed model is not automatically hypothetical. An intention-to-treat label does not prove treatment-policy handling. A per-protocol label does not fully specify what happened to every deviation, rescue treatment, or missing outcome.
Reviewers need the full chain: the stated clinical question, the intercurrent-event strategy, the data collected after the event, the analysis assumptions, and whether the estimator actually targets that question. Classifying a trial by one sentence in the abstract is often too optimistic.
A Practical Estimand-Aware Meta-Analysis Workflow
- Define the meta-analytic target before extraction. State the population, treatments, endpoint, time horizon, intercurrent-event strategy, and summary measure needed for the decision.
- Extract trial-level estimands, not just PICOs. Use protocols and statistical analysis plans when the publication is vague.
- Check analysis-to-estimand alignment. Record follow-up after discontinuation, rescue-medication handling, missing-data assumptions, and any post-randomization exclusions.
- Group compatible questions. Pool only when clinical and statistical compatibility are defensible. Separate syntheses may be more honest than one larger estimate.
- Treat strategy differences as a candidate source of heterogeneity. Explore them without claiming causation from a handful of trial-level contrasts.
- Report the evidence gap. If aggregate publications do not permit estimand alignment or re-estimation, say so. Missing detail is not evidence of compatibility.
When Harmonization Is Impossible
Participant-level data may allow analysts to derive more comparable estimands, but most reviews work from aggregate reports. One trial may stop collecting outcomes after discontinuation. Another may use retrieved dropout data. A third may impute under assumptions that are poorly described. In that setting, a clean re-analysis is often impossible.
Decision rule
Do not solve estimand incompatibility by renaming it “statistical heterogeneity.” First decide whether the effects belong to one question. Only then decide how to model their variation.
Sensitivity analyses can show how conclusions change under alternative classifications or restricted trial sets. They cannot reconstruct outcomes that were never collected, nor can they make a hypothetical effect answer a treatment-policy decision by statistical force.
Reviewer Red-Flag Checklist
Why This Matters for Aqrab
Evidence synthesis often looks rigorous at the calculation layer while the treatment question remains under-specified. The crucial clues are scattered across protocols, methods sections, missing-data notes, and supplements. A useful methodology critique reconnects them before judging the pooled result.
Use Aqrab Try to pressure-test whether a review's eligibility, extraction, and synthesis rules target one coherent effect. The valuable question is not merely “Can these estimates be combined?” It is “What decision would their combination actually inform?”
Methods Sources
- Remiro-Azócar A, Polavieja P, Boutmy E, et al. Incorporating estimands into meta-analyses of clinical trials. Research Synthesis Methods. Published online July 30, 2026. doi:10.1017/rsm.2026.10107.
- International Council for Harmonisation. ICH E9(R1): Addendum on estimands and sensitivity analysis in clinical trials. Step 4 guideline. 2019.
- European Medicines Agency. ICH E9 statistical principles for clinical trials: scientific guideline. Accessed August 31, 2026.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Hierarchical Testing in Clinical Trials: When a Significant Secondary Endpoint Is Still Descriptive
A practical hierarchical testing guide for clinical researchers. Reconstruct the prespecified testing path before treating a small p-value on a secondary endpoint as confirmatory evidence.
Randomized Withdrawal Trials: Why a Relapse-Prevention Win Is Not a New-Patient Effect
A practical randomized withdrawal trial guide for clinical researchers. Audit the run-in, responder enrichment, withdrawal contrast, safety, and target population before generalizing a maintenance-effect claim.
Allocation Concealment: The Randomized-Trial Safeguard That Works Before Assignment
A practical allocation concealment guide for clinical researchers and peer reviewers. Separate sequence generation, concealment, implementation, and blinding before trusting the word randomized.