Target Trial Emulation Cannot Randomize Clinical Judgment: A Pertussis Study Audit
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Target trial emulation is a disciplined way to state the experiment an observational study is trying to approximate. It is not a machine for turning clinical judgment into random allocation.
That distinction matters whenever clinicians choose treatment because of details that also predict the outcome. A new multicentre study comparing clarithromycin with azithromycin in hospitalised infants with pertussis is a useful case: the protocol is explicit, the question is important, and the remaining causal uncertainty is exactly what a methods review should make visible.
What the Study Reported
The retrospective cohort included 196 infants aged 12 months or younger across 11 Italian health centres. Clinicians selected clarithromycin or azithromycin according to clinical judgment. The analysis used propensity-score weighting to address measured confounding.
Treatment groups
146 received clarithromycin and 50 received azithromycin.
Primary outcome
Severe disease occurred in 48 and 30 infants, respectively.
Weighted result
Azithromycin was associated with higher odds of severe disease: OR 1.80, 95% CI 1.01–3.21.
The public abstract concludes that clarithromycin was associated with more favourable outcomes. That is appropriately association-shaped language. The reviewer's job is to determine how far the design can move the result toward a comparative treatment effect—and where it cannot.
The Blueprint Is Not the Assignment Mechanism
A target trial protocol aligns eligibility, treatment strategies, time zero, follow-up, outcomes, estimands, and analysis. This prevents avoidable design errors such as immortal time, inconsistent eligibility, or a vague comparator. Those gains are real.
But exchangeability remains a separate requirement. If a clinician chooses one macrolide for younger, frailer, more unstable, or differently presenting infants, the treatment groups can differ in prognosis before the first dose. Weighting can balance recorded variables. It cannot balance an undocumented bedside impression, an unrecorded referral history, or a center-specific preference that was never represented in the model.
The clean metaphor
A target trial is the blueprint for the experiment you wanted. Propensity weighting is the attempt to rebalance the materials you actually received. Neither redraws the history of why each clinician chose each treatment.
Interactive claim checker
How High Can the Causal Headline Climb?
Set each field to what the paper actually documents—not what the method label implies. The defaults mirror what can be established from the public abstract of the infant pertussis study: treatment by clinical judgment, with several diagnostics or endpoint details still requiring the full report.
Suggested claim ceiling
Keep the headline at a weighted association
The selected evidence leaves several plausible explanations besides treatment. Describe the association, show the diagnostics, and avoid language that makes emulation sound like randomization.
- • Clinical judgment can carry prognosis that the propensity model did not measure.
- • Without usable overlap and weight diagnostics, the estimate may depend on a small, atypical group.
- • The endpoint timeline is unclear, so post-treatment components remain an open audit question.
- • Thin events or wide uncertainty lower the ceiling for subgroup, safety, and transportability claims.
Teaching aid only. This is not a validated risk-of-bias instrument and it cannot certify causality. Its purpose is to keep the written claim no stronger than the design evidence a reader can verify.
Four Places the Causal Claim Can Still Break
| Audit gate | What to inspect | Why it changes the claim |
|---|---|---|
| Clinical-judgment confounding | Age in days, severity before treatment, symptom duration, referral pathway, contraindications, center preference | An unmeasured treatment driver can also predict deterioration. |
| Overlap and weights | Propensity distributions, extreme weights, effective sample size, trimming, balance after weighting | A 146-to-50 split may leave some infants without credible alternatives. |
| Endpoint timing | Every component of the Pertussis Severity Score, its measurement time, and its relationship to subsequent care | A post-treatment care process can become part of what the endpoint measures. |
| Sparse outcomes | Events behind secondary, safety, subgroup, and mortality estimates | Two deaths describe this cohort; they cannot carry a stable comparative mortality effect. |
Why Endpoint Definition Deserves Its Own Timeline
“Severe disease course” sounds clinically direct, but a score is a construction. The abstract identifies a Pertussis Severity Score above 5 as the primary outcome without listing its components. A full review should mark when each component becomes observable and whether treatment or subsequent care can affect it.
This does not prove endpoint contamination. It identifies the question the methods section must answer. A physiological state measured consistently after treatment is a legitimate outcome. A component driven by ICU admission, oxygen use, testing intensity, or another clinician response can partly reflect the care pathway. In a multicentre study, thresholds for those responses may also vary by hospital.
Draw the sequence: baseline presentation → antibiotic choice → evolving illness → care decisions → score assessment. Then decide whether the score estimates patient severity, a combined treatment-and-care strategy, or an inseparable mixture.
A Reviewer Checklist for Weighted Comparative Effectiveness
- Which pretreatment findings, clinician impressions, and care pathways made one antibiotic more likely than the other?
- Were those treatment drivers measured before the first dose, measured consistently across centers, and included in the design?
- Do propensity-score and weight distributions show usable support for both treatments in clinically comparable infants?
- What population remains after weighting, trimming, or stabilization, and is that population the one named in the conclusion?
- When was each primary-outcome component measured, and could treatment or subsequent care change the component itself?
- How many events support each secondary, safety, subgroup, and mortality claim—not only the primary contrast?
- Do alternative propensity models, truncation rules, center adjustments, and outcome definitions preserve the conclusion?
The fastest stress test is to ask what would make the two treatment groups exchange places. If one hospital strongly prefers clarithromycin, one clinician reserves azithromycin for a particular infant, or one pretreatment feature changes both drug choice and prognosis, that fact belongs in the design—not in a limitations paragraph after the estimate.
What the Estimate Can and Cannot Say
- It can report the observed, weighted contrast. The point estimate and interval describe the analysis that was performed.
- It cannot prove all relevant treatment drivers were measured. Covariate balance is limited to the covariates and representations shown.
- It should not turn two deaths into a mortality verdict. Rare events can be clinically important and statistically unstable at the same time.
- It does not automatically transport. Hospitalised Italian infants during one pertussis period may differ from infants in other systems, eras, or disease-severity ranges.
None of these cautions make the study useless. They make its contribution legible: an important, hypothesis-strengthening comparative effectiveness signal whose causal interpretation depends on measured treatment assignment, overlap, endpoint construction, and precision.
Where Aqrab Fits
Method labels are easy to reward. Design evidence is harder to inspect. Aqrab helps reviewers unpack a target-trial emulation into the exact questions that govern the claim: how treatment was chosen, what was known before time zero, whether both strategies had support, and whether the outcome sits downstream of treatment, care, or both.
If you are reviewing a real-world comparative effectiveness study, try Aqrab on the target-trial table and propensity-score methods together. For structured critique inside a research workflow, see the developer documentation.
Sources and Further Reading
- Lo Vecchio A, et al. Comparative effectiveness of clarithromycin and azithromycin for pertussis in hospitalised infants in Italy, the 2026 multicentre target-trial emulation used for this audit.
- Hernán MA, Robins JM. Using Big Data to Emulate a Target Trial, on specifying the target protocol and conditions needed for causal interpretation.
- Aqrab guide to target trial emulation, for the core protocol components.
- Aqrab guide to confounding by indication, for treatment assignment tied to prognosis.
- Aqrab guide to positivity and overlap, for weight and common-support diagnostics.
The Practical Bottom Line
Target trial emulation improves observational research by forcing the question into protocol form. That discipline is the beginning of causal credibility, not its completion.
When treatment follows clinical judgment, inspect what the clinician knew, what the dataset captured, where both strategies overlap, and when the outcome was created. The blueprint can clarify the experiment. It cannot randomize the past.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Vaccine Effectiveness Without Matching: Why Calendar Time Comes Before Pairing
A practical guide to calendar time in vaccine-effectiveness studies. Learn how changing uptake and infection hazards alter risk sets, estimands, and target-trial conclusions.
Comparator Selection in Observational Studies: Why the Control Group Changes the Question
A practical guide to comparator selection in observational studies. Learn how active comparators change the estimand, confounding structure, and interpretation of real-world evidence.
Treatment-Timing Effects: When “Earlier Is Better” Needs a Fair Clock
A practical guide to auditing earlier-is-better claims in observational studies. Test time zero, treatment strategies, evolving clinical decisions, positivity, and timing curves before reading a treatment gradient as biological.