← Back to Blog
Real-World EvidenceTarget Trial EmulationMethods Critique

Comparator Selection in Observational Studies: Why the Control Group Changes the Question

August 27, 2026·14 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Comparator selection in observational studies is not a clerical choice made after the exposure is named. It defines the treatment decision being studied, who can enter the comparison, and which reasons for treatment choice may still confound the result. Change the comparator and you have changed the trial.

That is why the same drug can appear strongly protective against one alternative and much less so against another. The discrepancy is not automatically an error. It is a prompt to ask whether the contrasts represent different causal questions, different matched populations, different comparator effects, or different degrees of residual confounding.

The Clean Metaphor: The Opponent Defines the Match

A runner's time means something different against a sprinter, a jogger, or an empty track.

The runner has not changed. The contest has. In comparative effectiveness research, the control group supplies the clinical alternative and the selection process against which the exposure is judged.

Interactive comparator audit

Change the Comparator, Change the Trial

Select the alternative treatment. The exposure stays the same, but the target question, matched population, bias structure, and interpretation change. Estimates below come from separate matched comparisons and should not be treated as a randomized head-to-head ranking.

Question being estimated

What happens if an eligible new user starts a GLP-1 receptor agonist rather than a DPP-4 inhibitor?

All-cause mortality

HR 0.61 (95% CI 0.52–0.71)

MACE

HR 0.83 (95% CI 0.71–0.98)

Coded heart failure

HR 0.71 (95% CI 0.60–0.85)

Interpretation

The mortality association looks large, but treatment channeling can remain even after matching. The authors caution that residual confounding is a likely explanation.

Reviewer move

Check whether the two drugs occupy the same line of care, why clinicians choose one over the other, and whether frailty, weight trajectory, kidney function, cost, and access were measured well enough.

Teaching note: the study used separate 1:1 propensity-score-matched cohorts. “Lower association reported” means the abstract states the direction but does not provide the usual-care point estimate.

A 2026 GLP-1 Study Makes the Problem Visible

Chen and colleagues used US electronic health records to emulate target trials among adults with type 2 diabetes and body mass index below 27 kg/m². New users of non-tirzepatide GLP-1 receptor agonists were compared separately with new users of DPP-4 inhibitors, new users of SGLT2 inhibitors, and usual care. Each comparison used its own 1:1 propensity-score matching, with follow-up up to 24 months.

Against DPP-4 inhibitors, GLP-1 initiation was associated with substantially lower all-cause mortality. Against SGLT2 inhibitors, the mortality estimate moved closer to no difference and its confidence interval included no effect. The authors did not turn that contrast into a marketing puzzle. They explicitly cautioned that the large mortality associations against DPP-4 inhibitors and usual care likely reflected residual confounding rather than a benefit of that magnitude.

The scientifically useful result is the instability

Comparator sensitivity can expose how much a conclusion depends on the alternative treatment and its selection pathway. It does not, by itself, reveal which estimate is unbiased.

The Comparator Changes Four Things at Once

What changesReviewer questionFailure if ignored
Clinical decisionWhich realistic alternatives could this patient receive now?An estimate that answers no recognizable care decision.
Target populationWho is jointly eligible for both strategies?Different populations mistaken for one treatment effect.
Confounding structureWhy does a clinician choose one option rather than the other?Treatment channeling survives adjustment.
Reference effectDoes the comparator itself affect the outcome?A relative contrast is misread as an absolute drug property.

“Treated versus untreated” often maximizes sample size while weakening exchangeability. Untreated patients may be untreated because of contraindications, frailty, limited access, lower engagement, mild disease, advanced disease, or a clinical plan that the data cannot observe. An active comparator can narrow those differences by placing both groups at a shared treatment decision. It cannot guarantee that the reasons for choosing between the active options were measured.

Why Propensity-Score Balance Is Necessary but Not Decisive

Matching can balance variables that were measured, represented correctly, and included in the score. A polished balance table says nothing direct about unrecorded frailty, clinician judgment, affordability, lifestyle treatment, contraindications hidden in free text, or incomplete disease severity. It also does not repair a comparator that occupies a different line of care.

The first audit question should therefore come before the standardized differences: could each patient plausibly have received either strategy at the same time? If the answer is no for a meaningful fraction of the cohort, the analysis may have limited positivity or be comparing different decision pathways. Statistical adjustment cannot manufacture a clinical choice that did not exist.

Do Not Rank Separate Hazard Ratios Like Trial Arms

The DPP-4, SGLT2, and usual-care analyses in the example were separate matched comparisons. Their GLP-1 groups were not necessarily identical, and each match can target a different overlap population. A hazard ratio of 0.61 in one cohort and 0.89 in another is therefore not a valid indirect estimate of the effect of DPP-4 inhibitors versus SGLT2 inhibitors.

Nor does overlap of confidence intervals settle the matter. Review baseline distributions, eligibility, match attrition, calendar time, treatment definitions, follow-up, and absolute risks within each comparison. Treat cross-comparator consistency as triangulation, not as a substitute for a prespecified head-to-head design.

A Practical Comparator Decision Rule

  1. Write the decision before naming the control. “Start A versus start B now” is more auditable than “A users versus controls.”
  2. Require joint eligibility. Apply the same baseline criteria at the same time zero to people who could receive either strategy.
  3. Map treatment channeling. List clinical, behavioral, access, prescriber, and system factors that drive the choice and predict the outcome.
  4. Prefer a credible active comparator when the decision supports one. Shared indication and line of care can reduce confounding by indication and surveillance differences.
  5. Define usual care as a strategy, not a label. State which drugs, delays, monitoring, escalation, and crossovers it contains.
  6. Triangulate deliberately. Alternative comparators are sensitivity analyses only when their different estimands and populations are made explicit.

Reviewer Red-Flag Checklist

The comparator is described only as “controls,” “nonusers,” or “usual care.”
Eligibility is defined for the treated group but not applied symmetrically to the comparator group.
Treatment initiation and start of follow-up are aligned in one group but ambiguous in the other.
Propensity-score balance is shown, but clinical reasons for choosing each treatment are not discussed.
Several comparator analyses are presented as replications even though they target different decisions and populations.
The paper compares hazard-ratio magnitudes across separately matched cohorts as if they came from one randomized trial.
A dramatic result against nonuse is privileged over a weaker result against a credible active alternative.
Residual confounding appears only as a generic sentence rather than an outcome- and comparator-specific argument.

Comparator sensitivity is not automatically evidence of misconduct or analytic failure. It is evidence that the paper must explain which clinical question each contrast answers and why the remaining bias is plausibly small enough for that conclusion.

Why This Matters for Aqrab

Comparator logic is scattered across a paper. Eligibility sits in the methods, clinical channeling is partly visible in baseline tables, overlap appears in matching diagnostics, treatment crossover may be buried in follow-up definitions, and the final causal claim lands in the abstract. A credible critique reconnects those pieces before rewarding a balanced table.

Use Aqrab Try to pressure-test whether an observational treatment claim survives its comparator choice. The sharp question is not “What was the control group?” It is “What alternative decision did this group make identifiable?”

Methods Sources

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive