Treatment History Is Part of the Estimand: Why Starting and Switching Answer Different Questions
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Treatment history is part of the estimand, not a nuisance column to adjust away. A patient starting a first therapy faces a different decision from a stable patient considering a switch. Put both patients in one exposure group, and the analysis may produce a precise answer to a question no clinician ever asked.
This matters whenever observational data are used to compare medications, devices, or care strategies. “Users of treatment A versus users of treatment B” sounds tidy. Yet current use can hide years of response, intolerance, survival, discontinuation, and prior treatment selection. A credible target trial begins at a real decision point and keeps patients with comparable treatment histories in the same question.
The Clean Metaphor: Treatment History Is the Runway
The same aircraft cannot take off from two different runways and claim it made the same journey.
Treatment initiation launches from no prior therapy. Switching launches from an existing regimen, with previous response and tolerance already on board. The drug name may match, but the counterfactual choice, baseline risk, comparator, and follow-up clock do not.
Interactive estimand explorer
Change the History, Change the Question
Explore the intention-to-treat estimates reported in the 2026 RESPOND analysis. This is a methods teaching reconstruction, not a treatment recommendation or an independent reanalysis.
Target question
Start an INSTI-based regimen versus start another antiretroviral regimen?
Time zero: First antiretroviral treatment initiation
Started INSTI
2.21%
Adjusted cardiovascular risk at 6 years
Started other ART
1.18%
Adjusted cardiovascular risk at 6 years
Risk ratio
1.87
95% CI 1.06 to 3.26
Absolute risk difference
+1.03 points
95% CI 0.08 to 2.07
Methods reading
The six-year curves have separated, but this cohort contributed only 95 cardiovascular events overall.
Neither scale is the “real” effect by itself. Read the relative contrast, absolute burden, uncertainty, population, and horizon together.
Why Initiators and Switchers Are Different Target Trials
For treatment-naïve patients, the decision might be: start regimen A or start regimen B today? Time zero is first treatment initiation. Eligibility, covariates, and treatment assignment can all be aligned at that point. This is the familiar active-comparator new-user design.
For treatment-experienced patients, the decision is different: switch from the current regimen to A or remain on an eligible alternative? There may be no single natural baseline because the switching decision can recur. One solution is to emulate a sequence of trials—for example, one each month—reassessing eligibility, treatment history, confounders, and assignment at every trial start.
Pooling the two populations does not merely reduce detail. It mixes treatment strategies. Initiators and switchers can differ in age, disease control, prior response, duration of care, contraindications, and the reasons clinicians chose the current regimen. An adjustment model can compare measured covariates. It cannot turn two different clinical decisions into one estimand.
A Current Example: One Drug Class, Two Risk Patterns
A 2026 RESPOND cohort analysis emulated separate target trials for integrase inhibitor–based antiretroviral therapy. Among 7,111 treatment-naïve adults, it compared starting an integrase inhibitor regimen with starting other antiretroviral therapy. Among 22,921 treatment-experienced adults with viral suppression, it compared switching to an integrase inhibitor with remaining on an eligible integrase-inhibitor-free regimen.
The patterns differed. In the treatment-naïve analysis, the adjusted cardiovascular risk curves gradually separated, reaching a six-year risk ratio of 1.87 and an absolute risk difference of 1.03 percentage points. But only 95 cardiovascular events occurred, so the long-horizon estimate deserves caution. In treatment-experienced patients, the association was strongest early: a one-year risk ratio of 1.55 with an absolute difference of 0.26 points. By six years, the risk ratio was 0.93 and the risk difference was −0.22 points, both compatible with no clear difference.
This is not evidence that one estimate is correct and the other is wrong. They describe different populations, decision points, and horizons. The investigators also noted residual confounding, treatment channeling, differential healthcare contact, cohort-selection sensitivity, and possible differences within the drug class. The constructive reading is methodological: separating treatment histories made the safety signal more interpretable without making the observational design infallible.
Relative and Absolute Risk Answer Different Parts of the Question
A risk ratio describes proportional contrast. A risk difference describes how many additional or fewer events are estimated in the study population over a stated horizon. When baseline risk is low, a large ratio can coexist with a small absolute difference. Neither measure should be used to silence the other.
Relative scale
Useful for proportional association and cross-population patterning, but easy to overdramatize when events are uncommon.
Absolute scale
Useful for clinical burden and decisions, but tied to the population's baseline risk and follow-up horizon.
Always report both with uncertainty and the underlying risks. “A 55% increase” is incomplete without the one-year risks, the 0.26-point difference, and the confidence intervals. “Only 0.26 points” is also incomplete when evaluating a preventable safety outcome across a large population. Interpretation is not a contest between scales; it is the discipline of keeping scale, horizon, and target population attached.
A Five-Question Treatment-History Audit
| Audit question | What a credible study shows | Red flag |
|---|---|---|
| What decision is being emulated? | Initiate, switch, continue, stop, or intensify | “Use” versus “non-use” |
| Who shares the same history? | Comparable prior therapies and eligibility | Naïve and prevalent users pooled |
| Where is time zero? | Eligibility, assignment, and follow-up align | Baseline chosen after treatment history unfolds |
| Is the comparator a real alternative? | Same indication and decision point | A clinically mixed “other” bucket |
| How is effect presented? | Risks, ratio, difference, horizon, and intervals | One dramatic relative estimate |
Reviewer Red-Flag Checklist
- New users, long-term users, switchers, and restarters share one exposure category.
- Prior treatment duration, response, intolerance, and discontinuation are absent from eligibility.
- Time zero is a convenient database date rather than a treatment decision.
- Switchers are compared with everyone who did not switch, regardless of whether switching was plausible.
- Baseline covariates are measured after treatment initiation or after early response is known.
- A long follow-up estimate is emphasized despite sparse events and wide uncertainty.
- Relative effects are reported without adjusted risks or absolute differences.
- A class effect is claimed when one treatment dominates exposure or comparators differ materially.
- A target-trial label is treated as proof that residual confounding and selection are gone.
Why This Matters for Aqrab
Many methods sections name a propensity score, weighting scheme, or target trial while leaving the treatment decision vague. A useful critique works in the opposite order: identify the decision, reconstruct treatment history, align time zero, then judge whether the model supports that question. The estimand is built before it is estimated.
Use Aqrab Try to ask whether a study's eligibility, treatment history, comparator, clock, and effect scale support its claim. Teams building repeatable review workflows can explore the developer tools.
Methods Sources
- Surial B, et al. Cardiovascular Disease Risk with Integrase Inhibitor-based Antiretroviral Therapy: A Causal Inference Analysis in RESPOND. Open Forum Infectious Diseases. 2026;13(8):ofag490.
- Hernán MA, Robins JM. Using Big Data to Emulate a Target Trial When a Randomized Trial Is Not Available. American Journal of Epidemiology. 2016;183(8):758–764.
- Ray WA. Evaluating Medication Effects Outside of Clinical Trials: New-User Designs. American Journal of Epidemiology. 2003;158(9):915–920.
- Lund JL, Richardson DB, Stürmer T. The active comparator, new user study design in pharmacoepidemiology: historical foundations and contemporary application. Current Epidemiology Reports. 2015;2(4):221–228.
- Darzi AJ, et al. Interpreting results from randomized controlled trials: What measures to focus on in clinical practice. Eye. 2023;37(15):3055–3058.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Target Trial Emulation Cannot Randomize Clinical Judgment: A Pertussis Study Audit
A practical target-trial emulation audit using an infant pertussis study. Check clinical-judgment confounding, propensity-score overlap, endpoint timing, sparse outcomes, and claim strength.
Vaccine Effectiveness Without Matching: Why Calendar Time Comes Before Pairing
A practical guide to calendar time in vaccine-effectiveness studies. Learn how changing uptake and infection hazards alter risk sets, estimands, and target-trial conclusions.
Comparator Selection in Observational Studies: Why the Control Group Changes the Question
A practical guide to comparator selection in observational studies. Learn how active comparators change the estimand, confounding structure, and interpretation of real-world evidence.