← Back to Blog
Evidence SynthesisClinical TrialsMethods Critique

Estimands in Meta-Analysis: When a Shared PICO Still Pools Different Questions

August 31, 2026·14 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Estimands in meta-analysis matter because two trials can share the same PICO and still estimate different treatment effects. The population matches. The drugs match. The endpoint and follow-up match. Yet one trial asks what happened after assignment, including discontinuation and rescue therapy, while another asks what would have happened had those events never occurred.

Pool those estimates without noticing the difference and the calculation may be flawless while the question is incoherent. PICO aligns the nouns. The estimand aligns the treatment question.

The Clean Example: Rescue Therapy Changes the Question

Imagine four randomized trials comparing two glucose-lowering treatments over 52 weeks. In every trial, some participants discontinue treatment or start rescue medication. Those events happen after assignment, affect the observed outcome, and complicate what “the treatment effect” means.

Treatment-policy question

What is the effect of assignment to each treatment strategy, regardless of later discontinuation or rescue? Those downstream events remain part of the practical treatment experience.

Hypothetical question

What would the effect have been if discontinuation or rescue had not occurred? This targets a counterfactual efficacy contrast and usually requires additional assumptions.

Neither question is universally superior. They serve different decisions. The error is to combine them as though they were noisy measurements of one automatic truth.

Interactive pooling audit

Same PICO, Different Treatment Questions

These four simulated trials share the same population, treatments, endpoint, and 52-week horizon. Negative values favor the intervention on a clinical-score scale. Change the target estimand to see why PICO alignment alone can produce a polished answer to a mixed question.

Trial A

Treatment policy

-3.2 points

Follow-up continues after discontinuation or rescue therapy.

Included

Trial B

Hypothetical

-7.1 points

Models the outcome had discontinuation or rescue therapy not occurred.

Included

Trial C

Treatment policy

-4.0 points

Uses outcomes regardless of treatment discontinuation or rescue therapy.

Included

Trial D

Hypothetical

-7.7 points

Targets efficacy in a world without discontinuation or rescue therapy.

Included

Illustrative fixed-effect summary

-5.5 points

95% interval -6.4 points to -4.7 points · 4 trials

Estimands mixed

The arithmetic is valid; the target is muddy. This summary averages pragmatic effectiveness with efficacy in a hypothetical world without discontinuation or rescue. Statistical compatibility does not make those questions interchangeable.

Teaching note: the values are simulated and the fixed-effect summaries are deliberately simple. Real synthesis also needs a justified heterogeneity model, compatible effect measures, uncertainty about classification, and careful review of whether trial-level analyses actually estimate the stated estimands.

What an Estimand Adds Beyond PICO

PICO remains indispensable for review scope. It describes who, which interventions, which comparison, and which outcome. The estimand framework adds the precision needed to understand the effect each trial is trying to estimate.

AttributeExtraction questionPooling risk if vague
PopulationWhich patients does the effect describe?Different eligibility or post-randomization subsets
Treatment conditionsWhat exactly is being compared?Different versions, doses, or background care
Variable or endpointWhat outcome and time point define the contrast?Nominally similar but clinically different outcomes
Intercurrent-event strategiesHow are discontinuation, rescue, switching, or death reflected in the question?Effectiveness and hypothetical efficacy mixed together
Population-level summaryDifference in means, risk ratio, hazard ratio, or another summary?Different scales or summaries treated as equivalent

Five Strategies Are Not Five Analysis Methods

ICH E9(R1) describes treatment policy, hypothetical, composite, while-on-treatment, and principal-stratum strategies for reflecting intercurrent events in the clinical question. These labels describe the target effect, not a menu of statistical commands. A mixed model is not automatically hypothetical. An intention-to-treat label does not prove treatment-policy handling. A per-protocol label does not fully specify what happened to every deviation, rescue treatment, or missing outcome.

Reviewers need the full chain: the stated clinical question, the intercurrent-event strategy, the data collected after the event, the analysis assumptions, and whether the estimator actually targets that question. Classifying a trial by one sentence in the abstract is often too optimistic.

A Practical Estimand-Aware Meta-Analysis Workflow

  1. Define the meta-analytic target before extraction. State the population, treatments, endpoint, time horizon, intercurrent-event strategy, and summary measure needed for the decision.
  2. Extract trial-level estimands, not just PICOs. Use protocols and statistical analysis plans when the publication is vague.
  3. Check analysis-to-estimand alignment. Record follow-up after discontinuation, rescue-medication handling, missing-data assumptions, and any post-randomization exclusions.
  4. Group compatible questions. Pool only when clinical and statistical compatibility are defensible. Separate syntheses may be more honest than one larger estimate.
  5. Treat strategy differences as a candidate source of heterogeneity. Explore them without claiming causation from a handful of trial-level contrasts.
  6. Report the evidence gap. If aggregate publications do not permit estimand alignment or re-estimation, say so. Missing detail is not evidence of compatibility.

When Harmonization Is Impossible

Participant-level data may allow analysts to derive more comparable estimands, but most reviews work from aggregate reports. One trial may stop collecting outcomes after discontinuation. Another may use retrieved dropout data. A third may impute under assumptions that are poorly described. In that setting, a clean re-analysis is often impossible.

Decision rule

Do not solve estimand incompatibility by renaming it “statistical heterogeneity.” First decide whether the effects belong to one question. Only then decide how to model their variation.

Sensitivity analyses can show how conclusions change under alternative classifications or restricted trial sets. They cannot reconstruct outcomes that were never collected, nor can they make a hypothetical effect answer a treatment-policy decision by statistical force.

Reviewer Red-Flag Checklist

Eligibility requires the same PICO, but extraction never records treatment discontinuation, rescue therapy, switching, or death.
Trials using treatment-policy and hypothetical strategies are pooled without naming the meta-analytic target estimand.
A subgroup split by analysis label is treated as proof that the estimand strategy caused the heterogeneity.
The review infers an estimand from “intention to treat” or “per protocol” without checking the endpoint definition and missing-data method.
A pooled estimate is presented as real-world effectiveness even though most trials target hypothetical efficacy.
Unavailable participant-level data are quietly replaced by optimistic assumptions about estimand compatibility.

Why This Matters for Aqrab

Evidence synthesis often looks rigorous at the calculation layer while the treatment question remains under-specified. The crucial clues are scattered across protocols, methods sections, missing-data notes, and supplements. A useful methodology critique reconnects them before judging the pooled result.

Use Aqrab Try to pressure-test whether a review's eligibility, extraction, and synthesis rules target one coherent effect. The valuable question is not merely “Can these estimates be combined?” It is “What decision would their combination actually inform?”

Methods Sources

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive