Comparator Selection in Observational Studies: Why the Control Group Changes the Question
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Comparator selection in observational studies is not a clerical choice made after the exposure is named. It defines the treatment decision being studied, who can enter the comparison, and which reasons for treatment choice may still confound the result. Change the comparator and you have changed the trial.
That is why the same drug can appear strongly protective against one alternative and much less so against another. The discrepancy is not automatically an error. It is a prompt to ask whether the contrasts represent different causal questions, different matched populations, different comparator effects, or different degrees of residual confounding.
The Clean Metaphor: The Opponent Defines the Match
A runner's time means something different against a sprinter, a jogger, or an empty track.
The runner has not changed. The contest has. In comparative effectiveness research, the control group supplies the clinical alternative and the selection process against which the exposure is judged.
Interactive comparator audit
Change the Comparator, Change the Trial
Select the alternative treatment. The exposure stays the same, but the target question, matched population, bias structure, and interpretation change. Estimates below come from separate matched comparisons and should not be treated as a randomized head-to-head ranking.
Question being estimated
What happens if an eligible new user starts a GLP-1 receptor agonist rather than a DPP-4 inhibitor?
All-cause mortality
HR 0.61 (95% CI 0.52–0.71)
MACE
HR 0.83 (95% CI 0.71–0.98)
Coded heart failure
HR 0.71 (95% CI 0.60–0.85)
Interpretation
The mortality association looks large, but treatment channeling can remain even after matching. The authors caution that residual confounding is a likely explanation.
Reviewer move
Check whether the two drugs occupy the same line of care, why clinicians choose one over the other, and whether frailty, weight trajectory, kidney function, cost, and access were measured well enough.
Teaching note: the study used separate 1:1 propensity-score-matched cohorts. “Lower association reported” means the abstract states the direction but does not provide the usual-care point estimate.
A 2026 GLP-1 Study Makes the Problem Visible
Chen and colleagues used US electronic health records to emulate target trials among adults with type 2 diabetes and body mass index below 27 kg/m². New users of non-tirzepatide GLP-1 receptor agonists were compared separately with new users of DPP-4 inhibitors, new users of SGLT2 inhibitors, and usual care. Each comparison used its own 1:1 propensity-score matching, with follow-up up to 24 months.
Against DPP-4 inhibitors, GLP-1 initiation was associated with substantially lower all-cause mortality. Against SGLT2 inhibitors, the mortality estimate moved closer to no difference and its confidence interval included no effect. The authors did not turn that contrast into a marketing puzzle. They explicitly cautioned that the large mortality associations against DPP-4 inhibitors and usual care likely reflected residual confounding rather than a benefit of that magnitude.
The scientifically useful result is the instability
Comparator sensitivity can expose how much a conclusion depends on the alternative treatment and its selection pathway. It does not, by itself, reveal which estimate is unbiased.
The Comparator Changes Four Things at Once
| What changes | Reviewer question | Failure if ignored |
|---|---|---|
| Clinical decision | Which realistic alternatives could this patient receive now? | An estimate that answers no recognizable care decision. |
| Target population | Who is jointly eligible for both strategies? | Different populations mistaken for one treatment effect. |
| Confounding structure | Why does a clinician choose one option rather than the other? | Treatment channeling survives adjustment. |
| Reference effect | Does the comparator itself affect the outcome? | A relative contrast is misread as an absolute drug property. |
“Treated versus untreated” often maximizes sample size while weakening exchangeability. Untreated patients may be untreated because of contraindications, frailty, limited access, lower engagement, mild disease, advanced disease, or a clinical plan that the data cannot observe. An active comparator can narrow those differences by placing both groups at a shared treatment decision. It cannot guarantee that the reasons for choosing between the active options were measured.
Why Propensity-Score Balance Is Necessary but Not Decisive
Matching can balance variables that were measured, represented correctly, and included in the score. A polished balance table says nothing direct about unrecorded frailty, clinician judgment, affordability, lifestyle treatment, contraindications hidden in free text, or incomplete disease severity. It also does not repair a comparator that occupies a different line of care.
The first audit question should therefore come before the standardized differences: could each patient plausibly have received either strategy at the same time? If the answer is no for a meaningful fraction of the cohort, the analysis may have limited positivity or be comparing different decision pathways. Statistical adjustment cannot manufacture a clinical choice that did not exist.
Do Not Rank Separate Hazard Ratios Like Trial Arms
The DPP-4, SGLT2, and usual-care analyses in the example were separate matched comparisons. Their GLP-1 groups were not necessarily identical, and each match can target a different overlap population. A hazard ratio of 0.61 in one cohort and 0.89 in another is therefore not a valid indirect estimate of the effect of DPP-4 inhibitors versus SGLT2 inhibitors.
Nor does overlap of confidence intervals settle the matter. Review baseline distributions, eligibility, match attrition, calendar time, treatment definitions, follow-up, and absolute risks within each comparison. Treat cross-comparator consistency as triangulation, not as a substitute for a prespecified head-to-head design.
A Practical Comparator Decision Rule
- Write the decision before naming the control. “Start A versus start B now” is more auditable than “A users versus controls.”
- Require joint eligibility. Apply the same baseline criteria at the same time zero to people who could receive either strategy.
- Map treatment channeling. List clinical, behavioral, access, prescriber, and system factors that drive the choice and predict the outcome.
- Prefer a credible active comparator when the decision supports one. Shared indication and line of care can reduce confounding by indication and surveillance differences.
- Define usual care as a strategy, not a label. State which drugs, delays, monitoring, escalation, and crossovers it contains.
- Triangulate deliberately. Alternative comparators are sensitivity analyses only when their different estimands and populations are made explicit.
Reviewer Red-Flag Checklist
Comparator sensitivity is not automatically evidence of misconduct or analytic failure. It is evidence that the paper must explain which clinical question each contrast answers and why the remaining bias is plausibly small enough for that conclusion.
Why This Matters for Aqrab
Comparator logic is scattered across a paper. Eligibility sits in the methods, clinical channeling is partly visible in baseline tables, overlap appears in matching diagnostics, treatment crossover may be buried in follow-up definitions, and the final causal claim lands in the abstract. A credible critique reconnects those pieces before rewarding a balanced table.
Use Aqrab Try to pressure-test whether an observational treatment claim survives its comparator choice. The sharp question is not “What was the control group?” It is “What alternative decision did this group make identifiable?”
Methods Sources
- Chen SC, et al. Cardiorenal Mortality and Safety Outcomes of GLP-1 Receptor Agonists in Type 2 Diabetes With BMI Below 27 kg/m²: A Target Trial Emulation. Diabetes, Obesity and Metabolism. 2026. doi:10.1111/dom.71266.
- Sendor R, Stürmer T. Core concepts in pharmacoepidemiology: confounding by indication and the role of active comparators. Pharmacoepidemiology and Drug Safety. 2022;31(3):261–269.
- Hernán MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. American Journal of Epidemiology. 2016;183(8):758–764.
- Cashin AG, et al. Improving reporting of observational studies of interventions: the TARGET guideline. PLOS Medicine. 2025.
- National Institute for Health and Care Excellence. Methods for real-world studies of comparative effects. NICE real-world evidence framework.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Target Trial Emulation Cannot Randomize Clinical Judgment: A Pertussis Study Audit
A practical target-trial emulation audit using an infant pertussis study. Check clinical-judgment confounding, propensity-score overlap, endpoint timing, sparse outcomes, and claim strength.
Vaccine Effectiveness Without Matching: Why Calendar Time Comes Before Pairing
A practical guide to calendar time in vaccine-effectiveness studies. Learn how changing uptake and infection hazards alter risk sets, estimands, and target-trial conclusions.
Treatment-Timing Effects: When “Earlier Is Better” Needs a Fair Clock
A practical guide to auditing earlier-is-better claims in observational studies. Test time zero, treatment strategies, evolving clinical decisions, positivity, and timing curves before reading a treatment gradient as biological.