← Back to Blog
Target Trial EmulationReal-World EvidenceMethods Critique

How to Audit a Target-Trial Emulation: Four Decisions Behind One Mortality Estimate

August 15, 2026·14 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A target-trial emulation can look like a simple comparison: treatment A versus treatment B, followed by a mortality estimate. But the estimate is assembled from protocol decisions. Change the day patients enter, the adherence rule, the handling of transplantation, or the facilities contributing each treatment, and you may change the question before changing a single model coefficient.

A new multinational dialysis study makes that assembly visible. It compared hemodiafiltration with high-flux hemodialysis using registry data, began follow-up 91 days after dialysis initiation, defined sustained treatment as receiving the assigned modality for at least 90% of sessions, used inverse-probability weighting, and treated kidney transplantation as a competing event. This is exactly the kind of study that deserves more than checking whether “target trial emulation” appears in the title.

The Result Is Not the Question

The study reported an adjusted hazard ratio of 0.72 for all-cause mortality. At two years, the weighted cumulative incidence of death was 20.6% with hemodiafiltration and 22.3% with high-flux hemodialysis—an absolute difference of 1.7 percentage points.

The clean metaphor

The effect estimate is the address on the envelope. The protocol is the letter. If you read only the number, you still do not know what claim was delivered.

A hazard ratio of 0.72 is not the same statement as “28% fewer patients died by two years.” The hazard ratio is a relative rate comparison over follow-up under model assumptions. The two-year cumulative-incidence contrast is an absolute risk comparison at a specific horizon. Both can be useful. They answer on different scales and should not be blended into one dramatic sentence.

Interactive estimand translator

Turn the headline back into a protocol question

The defaults mirror key choices reported in the dialysis study abstract. Change one choice at a time and watch the population, treatment strategy, and assumptions move with it.

1. When does the study population enter?
2. What is the treatment strategy?
3. How is transplantation handled?
4. Can modality be separated from center?

The question you built

Among adults who are alive, observable, and still eligible 91 days after dialysis initiation, what would happen under strategies of sustaining each assigned modality for at least 90% of sessions, for death in a framework where transplantation changes the risk set as a competing event?

What the short headline hides

  • The estimate is conditional on reaching day 91. Early deaths, early transplants, and people no longer observable are outside this target population.
  • The 90% threshold defines the intervention. Credibility depends on measuring time-varying reasons for switching, missed sessions, and discontinuation—not merely labeling adherers.
  • The paper must state the competing-risk estimand. A cause-specific hazard, a subdistribution hazard, and a cumulative-incidence risk are not interchangeable.
  • Country or facility adjustment helps only where both modalities have credible within-center support. Review overlap and center-level treatment preference.

Teaching aid only. It translates design choices; it does not recalculate the study estimate or judge the clinical effectiveness of either dialysis modality.

Decision 1: A Day-91 Landmark Changes the Population

Beginning follow-up 91 days after dialysis initiation can create a clean, clinically recognizable baseline. It may allow treatment patterns and covariates to stabilize. It also means the estimate is not about every patient who starts dialysis.

To be eligible at day 91, a patient must survive, remain observable, and continue to meet the study's criteria until that point. Early deaths and early transitions occur before the analysis begins. That is not automatically bias; a landmark can be a legitimate design choice. The error is carrying the result back to dialysis initiation as though the first 90 days were represented.

Reviewer red flag

The conclusion says “among patients initiating dialysis” when the analytic population contains only those eligible and observed at a later landmark.

Decision 2: “At Least 90% of Sessions” Is the Intervention

A sustained strategy is not a baseline treatment label with better manners. Requiring a modality for at least 90% of sessions defines what it means to follow that strategy. A patient receiving it for 89% of sessions and one receiving it for 91% are placed on opposite sides of a threshold, even if their clinical experience is nearly identical.

The threshold may be clinically defensible, but reviewers should ask when adherence is assessed, which missed or switched sessions count, whether the rule uses future information, and how deviations are handled. If worsening health, vascular access, hospitalization, intolerance, or center capacity predicts both modality changes and mortality, then protocol adherence is prognostic. Excluding deviators creates selection; modeling only baseline treatment does not estimate sustained receipt.

A defensible analysis therefore needs time-updated information on reasons for deviation, adequate support for staying on each strategy, transparent weight diagnostics when weighting is used, and sensitivity analyses that change the adherence definition. The goal is not to find the threshold that produces the strongest effect. It is to show whether the clinical conclusion survives reasonable versions of the strategy.

Decision 3: Transplantation Changes the Outcome Question

Kidney transplantation is not ordinary administrative loss to follow-up. It is related to prognosis, access, eligibility, and treatment history, and it changes the clinical pathway. Calling it a competing event is only the start; the paper must still name the estimand.

Analysis quantityWhat it emphasizesReviewer question
Cause-specific hazardInstantaneous death rate among those still at riskDoes the interpretation stay on the hazard scale?
Cumulative incidenceAbsolute probability of death over time in the presence of transplantationIs the time horizon explicit?
Censor at transplantA world requiring assumptions about post-transplant censoringWhy is censoring plausibly noninformative, or how was it adjusted?

The key judgment is alignment: the outcome definition, competing-event method, summary measure, and clinical conclusion must describe the same question.

Decision 4: A Modality Can Travel With Its Center

In multinational registry data, treatment is delivered inside facilities. Hemodiafiltration availability, delivered volume, staffing, vascular-access practices, patient selection, monitoring, and clinical expertise can cluster by center. If one modality is concentrated in better-resourced or more experienced facilities, a pooled patient-level comparison can partly compare care systems.

Adding country or center to a model does not manufacture overlap. Reviewers need to know whether both modalities are actually used within the same facilities and patient strata, how concentrated treatment preference is, whether standard errors respect clustering, and whether within-center or center-restricted analyses support the result. The study authors appropriately caution that an apparent convective-volume gradient may reflect patient stability and center expertise rather than a causal dose–response relationship. The same instinct—separating treatment from its delivery context—is worth applying to the main contrast.

What the Sensitivity Analyses Can—and Cannot—Do

The reported results were broadly consistent across analyses addressing country, competing risks, informative censoring, protocol adherence, and an initiation-like exposure definition. Consistency across reasonable specifications is reassuring because it shows the headline is not tied to one visible switch.

It does not prove exchangeability, eliminate unmeasured differences in patient stability or facility expertise, or make every estimand identical. Sensitivity analyses are strongest when each one names the threat it probes, shows the estimate on the same interpretable scale, and explains why the alternative specification is plausible.

A Six-Question Reviewer Checklist

  1. Who reaches time zero? State who is excluded by any landmark or required baseline history.
  2. What treatment strategy could a clinician implement? Include timing, session threshold, switching, interruption, and allowable deviations.
  3. Does adherence use future information? Verify that classification and follow-up do not quietly grant immortal person-time.
  4. What predicts deviation? Demand time-varying covariates, censoring logic, weight distributions, truncation choices, and positivity diagnostics where relevant.
  5. What happens at transplantation? Match the competing-event method to the estimand and the language of the conclusion.
  6. Can treatment be separated from center? Inspect within-facility use, clustering, preference, expertise, and sensitivity to facility restriction or adjustment.

Why This Matters for Aqrab

Aqrab's credibility lives in the distance between spotting a method name and understanding the claim it permits. A target-trial emulation is not rigorous because it contains a protocol table, propensity weights, and a causal adjective. It is rigorous when the target population, treatment strategy, time zero, outcome, estimand, and identifying assumptions remain aligned from methods to abstract.

Use Aqrab Try to pressure-test whether an observational study's headline still matches the trial it actually emulated. The useful critique is not “this was not randomized.” It is the more precise sentence: “This estimate applies to these patients, under this strategy, over this horizon, if these measured processes were sufficient.”

Methods Anchors

The applied details come from Strippoli and colleagues' multinational dialysis target-trial emulation in the Journal of the American Society of Nephrology. The four-targets framework connects the target estimand, target population, target trial, and target validity. Work on emulating sustained treatment strategies with observational data shows why per-protocol effects require explicit deviation rules and adjustment for time-varying predictors of adherence. A practical target-trial emulation overview explains how the causal question determines whether an active-comparator, sequential-trial, or clone–censor–weight design is appropriate.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive