How to Audit a Target-Trial Emulation: Four Decisions Behind One Mortality Estimate
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
A target-trial emulation can look like a simple comparison: treatment A versus treatment B, followed by a mortality estimate. But the estimate is assembled from protocol decisions. Change the day patients enter, the adherence rule, the handling of transplantation, or the facilities contributing each treatment, and you may change the question before changing a single model coefficient.
A new multinational dialysis study makes that assembly visible. It compared hemodiafiltration with high-flux hemodialysis using registry data, began follow-up 91 days after dialysis initiation, defined sustained treatment as receiving the assigned modality for at least 90% of sessions, used inverse-probability weighting, and treated kidney transplantation as a competing event. This is exactly the kind of study that deserves more than checking whether “target trial emulation” appears in the title.
The Result Is Not the Question
The study reported an adjusted hazard ratio of 0.72 for all-cause mortality. At two years, the weighted cumulative incidence of death was 20.6% with hemodiafiltration and 22.3% with high-flux hemodialysis—an absolute difference of 1.7 percentage points.
The clean metaphor
The effect estimate is the address on the envelope. The protocol is the letter. If you read only the number, you still do not know what claim was delivered.
A hazard ratio of 0.72 is not the same statement as “28% fewer patients died by two years.” The hazard ratio is a relative rate comparison over follow-up under model assumptions. The two-year cumulative-incidence contrast is an absolute risk comparison at a specific horizon. Both can be useful. They answer on different scales and should not be blended into one dramatic sentence.
Interactive estimand translator
Turn the headline back into a protocol question
The defaults mirror key choices reported in the dialysis study abstract. Change one choice at a time and watch the population, treatment strategy, and assumptions move with it.
The question you built
Among adults who are alive, observable, and still eligible 91 days after dialysis initiation, what would happen under strategies of sustaining each assigned modality for at least 90% of sessions, for death in a framework where transplantation changes the risk set as a competing event?
What the short headline hides
- • The estimate is conditional on reaching day 91. Early deaths, early transplants, and people no longer observable are outside this target population.
- • The 90% threshold defines the intervention. Credibility depends on measuring time-varying reasons for switching, missed sessions, and discontinuation—not merely labeling adherers.
- • The paper must state the competing-risk estimand. A cause-specific hazard, a subdistribution hazard, and a cumulative-incidence risk are not interchangeable.
- • Country or facility adjustment helps only where both modalities have credible within-center support. Review overlap and center-level treatment preference.
Teaching aid only. It translates design choices; it does not recalculate the study estimate or judge the clinical effectiveness of either dialysis modality.
Decision 1: A Day-91 Landmark Changes the Population
Beginning follow-up 91 days after dialysis initiation can create a clean, clinically recognizable baseline. It may allow treatment patterns and covariates to stabilize. It also means the estimate is not about every patient who starts dialysis.
To be eligible at day 91, a patient must survive, remain observable, and continue to meet the study's criteria until that point. Early deaths and early transitions occur before the analysis begins. That is not automatically bias; a landmark can be a legitimate design choice. The error is carrying the result back to dialysis initiation as though the first 90 days were represented.
Reviewer red flag
The conclusion says “among patients initiating dialysis” when the analytic population contains only those eligible and observed at a later landmark.
Decision 2: “At Least 90% of Sessions” Is the Intervention
A sustained strategy is not a baseline treatment label with better manners. Requiring a modality for at least 90% of sessions defines what it means to follow that strategy. A patient receiving it for 89% of sessions and one receiving it for 91% are placed on opposite sides of a threshold, even if their clinical experience is nearly identical.
The threshold may be clinically defensible, but reviewers should ask when adherence is assessed, which missed or switched sessions count, whether the rule uses future information, and how deviations are handled. If worsening health, vascular access, hospitalization, intolerance, or center capacity predicts both modality changes and mortality, then protocol adherence is prognostic. Excluding deviators creates selection; modeling only baseline treatment does not estimate sustained receipt.
A defensible analysis therefore needs time-updated information on reasons for deviation, adequate support for staying on each strategy, transparent weight diagnostics when weighting is used, and sensitivity analyses that change the adherence definition. The goal is not to find the threshold that produces the strongest effect. It is to show whether the clinical conclusion survives reasonable versions of the strategy.
Decision 3: Transplantation Changes the Outcome Question
Kidney transplantation is not ordinary administrative loss to follow-up. It is related to prognosis, access, eligibility, and treatment history, and it changes the clinical pathway. Calling it a competing event is only the start; the paper must still name the estimand.
| Analysis quantity | What it emphasizes | Reviewer question |
|---|---|---|
| Cause-specific hazard | Instantaneous death rate among those still at risk | Does the interpretation stay on the hazard scale? |
| Cumulative incidence | Absolute probability of death over time in the presence of transplantation | Is the time horizon explicit? |
| Censor at transplant | A world requiring assumptions about post-transplant censoring | Why is censoring plausibly noninformative, or how was it adjusted? |
The key judgment is alignment: the outcome definition, competing-event method, summary measure, and clinical conclusion must describe the same question.
Decision 4: A Modality Can Travel With Its Center
In multinational registry data, treatment is delivered inside facilities. Hemodiafiltration availability, delivered volume, staffing, vascular-access practices, patient selection, monitoring, and clinical expertise can cluster by center. If one modality is concentrated in better-resourced or more experienced facilities, a pooled patient-level comparison can partly compare care systems.
Adding country or center to a model does not manufacture overlap. Reviewers need to know whether both modalities are actually used within the same facilities and patient strata, how concentrated treatment preference is, whether standard errors respect clustering, and whether within-center or center-restricted analyses support the result. The study authors appropriately caution that an apparent convective-volume gradient may reflect patient stability and center expertise rather than a causal dose–response relationship. The same instinct—separating treatment from its delivery context—is worth applying to the main contrast.
What the Sensitivity Analyses Can—and Cannot—Do
The reported results were broadly consistent across analyses addressing country, competing risks, informative censoring, protocol adherence, and an initiation-like exposure definition. Consistency across reasonable specifications is reassuring because it shows the headline is not tied to one visible switch.
It does not prove exchangeability, eliminate unmeasured differences in patient stability or facility expertise, or make every estimand identical. Sensitivity analyses are strongest when each one names the threat it probes, shows the estimate on the same interpretable scale, and explains why the alternative specification is plausible.
A Six-Question Reviewer Checklist
- Who reaches time zero? State who is excluded by any landmark or required baseline history.
- What treatment strategy could a clinician implement? Include timing, session threshold, switching, interruption, and allowable deviations.
- Does adherence use future information? Verify that classification and follow-up do not quietly grant immortal person-time.
- What predicts deviation? Demand time-varying covariates, censoring logic, weight distributions, truncation choices, and positivity diagnostics where relevant.
- What happens at transplantation? Match the competing-event method to the estimand and the language of the conclusion.
- Can treatment be separated from center? Inspect within-facility use, clustering, preference, expertise, and sensitivity to facility restriction or adjustment.
Why This Matters for Aqrab
Aqrab's credibility lives in the distance between spotting a method name and understanding the claim it permits. A target-trial emulation is not rigorous because it contains a protocol table, propensity weights, and a causal adjective. It is rigorous when the target population, treatment strategy, time zero, outcome, estimand, and identifying assumptions remain aligned from methods to abstract.
Use Aqrab Try to pressure-test whether an observational study's headline still matches the trial it actually emulated. The useful critique is not “this was not randomized.” It is the more precise sentence: “This estimate applies to these patients, under this strategy, over this horizon, if these measured processes were sufficient.”
Methods Anchors
The applied details come from Strippoli and colleagues' multinational dialysis target-trial emulation in the Journal of the American Society of Nephrology. The four-targets framework connects the target estimand, target population, target trial, and target validity. Work on emulating sustained treatment strategies with observational data shows why per-protocol effects require explicit deviation rules and adjustment for time-varying predictors of adherence. A practical target-trial emulation overview explains how the causal question determines whether an active-comparator, sequential-trial, or clone–censor–weight design is appropriate.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Target Trial Emulation Cannot Randomize Clinical Judgment: A Pertussis Study Audit
A practical target-trial emulation audit using an infant pertussis study. Check clinical-judgment confounding, propensity-score overlap, endpoint timing, sparse outcomes, and claim strength.
Vaccine Effectiveness Without Matching: Why Calendar Time Comes Before Pairing
A practical guide to calendar time in vaccine-effectiveness studies. Learn how changing uptake and infection hazards alter risk sets, estimands, and target-trial conclusions.
Comparator Selection in Observational Studies: Why the Control Group Changes the Question
A practical guide to comparator selection in observational studies. Learn how active comparators change the estimand, confounding structure, and interpretation of real-world evidence.