← Back to Blog
Target Trial EmulationReal-World EvidenceMethods Critique

Target Trial Emulation Cannot Randomize Clinical Judgment: A Pertussis Study Audit

September 5, 2026·14 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Target trial emulation is a disciplined way to state the experiment an observational study is trying to approximate. It is not a machine for turning clinical judgment into random allocation.

That distinction matters whenever clinicians choose treatment because of details that also predict the outcome. A new multicentre study comparing clarithromycin with azithromycin in hospitalised infants with pertussis is a useful case: the protocol is explicit, the question is important, and the remaining causal uncertainty is exactly what a methods review should make visible.

What the Study Reported

The retrospective cohort included 196 infants aged 12 months or younger across 11 Italian health centres. Clinicians selected clarithromycin or azithromycin according to clinical judgment. The analysis used propensity-score weighting to address measured confounding.

Treatment groups

146 received clarithromycin and 50 received azithromycin.

Primary outcome

Severe disease occurred in 48 and 30 infants, respectively.

Weighted result

Azithromycin was associated with higher odds of severe disease: OR 1.80, 95% CI 1.01–3.21.

The public abstract concludes that clarithromycin was associated with more favourable outcomes. That is appropriately association-shaped language. The reviewer's job is to determine how far the design can move the result toward a comparative treatment effect—and where it cannot.

The Blueprint Is Not the Assignment Mechanism

A target trial protocol aligns eligibility, treatment strategies, time zero, follow-up, outcomes, estimands, and analysis. This prevents avoidable design errors such as immortal time, inconsistent eligibility, or a vague comparator. Those gains are real.

But exchangeability remains a separate requirement. If a clinician chooses one macrolide for younger, frailer, more unstable, or differently presenting infants, the treatment groups can differ in prognosis before the first dose. Weighting can balance recorded variables. It cannot balance an undocumented bedside impression, an unrecorded referral history, or a center-specific preference that was never represented in the model.

The clean metaphor

A target trial is the blueprint for the experiment you wanted. Propensity weighting is the attempt to rebalance the materials you actually received. Neither redraws the history of why each clinician chose each treatment.

Interactive claim checker

How High Can the Causal Headline Climb?

Set each field to what the paper actually documents—not what the method label implies. The defaults mirror what can be established from the public abstract of the infant pertussis study: treatment by clinical judgment, with several diagnostics or endpoint details still requiring the full report.

Suggested claim ceiling

Keep the headline at a weighted association

The selected evidence leaves several plausible explanations besides treatment. Describe the association, show the diagnostics, and avoid language that makes emulation sound like randomization.

  • Clinical judgment can carry prognosis that the propensity model did not measure.
  • Without usable overlap and weight diagnostics, the estimate may depend on a small, atypical group.
  • The endpoint timeline is unclear, so post-treatment components remain an open audit question.
  • Thin events or wide uncertainty lower the ceiling for subgroup, safety, and transportability claims.

Teaching aid only. This is not a validated risk-of-bias instrument and it cannot certify causality. Its purpose is to keep the written claim no stronger than the design evidence a reader can verify.

Four Places the Causal Claim Can Still Break

Audit gateWhat to inspectWhy it changes the claim
Clinical-judgment confoundingAge in days, severity before treatment, symptom duration, referral pathway, contraindications, center preferenceAn unmeasured treatment driver can also predict deterioration.
Overlap and weightsPropensity distributions, extreme weights, effective sample size, trimming, balance after weightingA 146-to-50 split may leave some infants without credible alternatives.
Endpoint timingEvery component of the Pertussis Severity Score, its measurement time, and its relationship to subsequent careA post-treatment care process can become part of what the endpoint measures.
Sparse outcomesEvents behind secondary, safety, subgroup, and mortality estimatesTwo deaths describe this cohort; they cannot carry a stable comparative mortality effect.

Why Endpoint Definition Deserves Its Own Timeline

“Severe disease course” sounds clinically direct, but a score is a construction. The abstract identifies a Pertussis Severity Score above 5 as the primary outcome without listing its components. A full review should mark when each component becomes observable and whether treatment or subsequent care can affect it.

This does not prove endpoint contamination. It identifies the question the methods section must answer. A physiological state measured consistently after treatment is a legitimate outcome. A component driven by ICU admission, oxygen use, testing intensity, or another clinician response can partly reflect the care pathway. In a multicentre study, thresholds for those responses may also vary by hospital.

Draw the sequence: baseline presentation → antibiotic choice → evolving illness → care decisions → score assessment. Then decide whether the score estimates patient severity, a combined treatment-and-care strategy, or an inseparable mixture.

A Reviewer Checklist for Weighted Comparative Effectiveness

  • Which pretreatment findings, clinician impressions, and care pathways made one antibiotic more likely than the other?
  • Were those treatment drivers measured before the first dose, measured consistently across centers, and included in the design?
  • Do propensity-score and weight distributions show usable support for both treatments in clinically comparable infants?
  • What population remains after weighting, trimming, or stabilization, and is that population the one named in the conclusion?
  • When was each primary-outcome component measured, and could treatment or subsequent care change the component itself?
  • How many events support each secondary, safety, subgroup, and mortality claim—not only the primary contrast?
  • Do alternative propensity models, truncation rules, center adjustments, and outcome definitions preserve the conclusion?

The fastest stress test is to ask what would make the two treatment groups exchange places. If one hospital strongly prefers clarithromycin, one clinician reserves azithromycin for a particular infant, or one pretreatment feature changes both drug choice and prognosis, that fact belongs in the design—not in a limitations paragraph after the estimate.

What the Estimate Can and Cannot Say

  1. It can report the observed, weighted contrast. The point estimate and interval describe the analysis that was performed.
  2. It cannot prove all relevant treatment drivers were measured. Covariate balance is limited to the covariates and representations shown.
  3. It should not turn two deaths into a mortality verdict. Rare events can be clinically important and statistically unstable at the same time.
  4. It does not automatically transport. Hospitalised Italian infants during one pertussis period may differ from infants in other systems, eras, or disease-severity ranges.

None of these cautions make the study useless. They make its contribution legible: an important, hypothesis-strengthening comparative effectiveness signal whose causal interpretation depends on measured treatment assignment, overlap, endpoint construction, and precision.

Where Aqrab Fits

Method labels are easy to reward. Design evidence is harder to inspect. Aqrab helps reviewers unpack a target-trial emulation into the exact questions that govern the claim: how treatment was chosen, what was known before time zero, whether both strategies had support, and whether the outcome sits downstream of treatment, care, or both.

If you are reviewing a real-world comparative effectiveness study, try Aqrab on the target-trial table and propensity-score methods together. For structured critique inside a research workflow, see the developer documentation.

Sources and Further Reading

The Practical Bottom Line

Target trial emulation improves observational research by forcing the question into protocol form. That discipline is the beginning of causal credibility, not its completion.

When treatment follows clinical judgment, inspect what the clinician knew, what the dataset captured, where both strategies overlap, and when the outcome was created. The blueprint can clarify the experiment. It cannot randomize the past.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive