Desirability of Outcome Ranking: When Benefit–Risk Depends on Who Ranks the Outcomes
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Desirability of outcome ranking (DOOR) asks a patient-level benefit–risk question: if we draw one patient from each treatment group, what is the probability that the patient receiving the new strategy has the more desirable overall outcome? That can be more clinically honest than reporting efficacy in one table and toxicity in another, as though benefits and harms happened to different people.
But DOOR does not make value judgments disappear. It makes them operational. Someone must decide whether recovery with serious toxicity ranks above non-recovery without it, whether death always ranks last, and how ties are handled. A polished probability is only as credible as that ordering.
Why Separate Endpoint Tables Can Miss the Patient
Trials often choose one primary efficacy endpoint, then list adverse events as secondary outcomes. This protects a clear hypothesis, but it can fragment the clinical story. A treatment may improve recovery while causing more serious toxicity. Marginal summaries tell us how often each event occurred; they may not show which combinations occurred within the same patient.
A broad composite endpoint creates a different problem. “Death, treatment failure, or serious toxicity” may gain events and power, but an ordinary time-to-first-event analysis can treat components with very different clinical importance as equivalent. It can also discard what happens after the first component.
DOOR starts by placing each patient into an ordered overall-outcome category that combines relevant benefits and harms. The analysis then compares the entire distribution between treatment strategies. The attraction is not that multiple outcomes become simple. It is that their trade-off becomes explicit enough to inspect.
The clean metaphor
DOOR is a staircase, not a blender. Patients occupy ordered steps. The method compares where groups land, but clinical judgment still decides which step is above which.
Interactive ranking-sensitivity explorer
Change one clinical priority. Watch the conclusion move.
This teaching dataset compares 100 patients assigned to a shorter antibiotic course with 100 assigned to a longer course. The only disputed judgment is whether recovery with serious toxicity is preferable to non-recovery without serious toxicity.
| Patient outcome | Short course | Long course |
|---|---|---|
| Recovered, no serious toxicity | 36 | 38 |
| Recovered, with serious toxicity | 12 | 24 |
| Not recovered, no serious toxicity | 34 | 22 |
| Not recovered, with serious toxicity | 8 | 7 |
| Death | 10 | 9 |
Which outcome should rank higher?
DOOR probability
45.8%
A randomly selected short-course patient has a 45.8% probability of a better outcome than a randomly selected long-course patient, with ties split equally. This ranking slightly favors the long course.
Teaching simplification: these counts are fictional and the switch is intentional. A real analysis should prespecify clinically defensible ranks, show component outcomes, quantify uncertainty, and test plausible alternative rankings rather than searching for the most favorable one.
How to Read a DOOR Probability
The usual summary is the probability that a randomly selected patient assigned to one strategy has a more desirable outcome than a randomly selected patient assigned to the comparator, with half credit for ties. A probability of 50% indicates no tendency for either group to rank higher. Values above 50% favor the named strategy; values below 50% favor the comparator.
This is a probabilistic index, not a risk difference. A DOOR probability of 56% does not mean 56% of patients benefit, nor does it mean the strategy improves outcomes by six percentage points. It describes pairwise ordering across the two outcome distributions. Report the confidence interval and the category-specific results that produced it.
| Summary | What it preserves | What it can hide |
|---|---|---|
| Separate outcomes | Effect on each benefit and harm | How outcomes combine within patients |
| Time to first composite event | Event timing under a defined composite | Component importance and later events |
| DOOR probability | Whole-patient ordinal ranking | Component effects if only one number is shown |
Four Ways a Sensible Ranking Becomes a Fragile Result
1. The ranks are written after the data are seen
Changing the category order until the probability becomes favorable is outcome switching in ordinal clothing. Define the outcomes, their timing, and their ordering before unblinded comparison.
2. Clinicians rank outcomes on behalf of every patient
Some orderings are clear; others are preference-sensitive. A patient may accept severe transient toxicity for recovery, while another may not. Use patient input where trade-offs are contestable and show plausible alternative rankings.
3. One probability replaces the component outcomes
Two treatments can have similar DOOR probabilities for very different reasons. Always show the distribution across categories and the underlying benefit and harm components. The global summary should organize the evidence, not erase it.
4. Ordinal categories are treated as equal distances
Ranking says one category is preferable to another; it does not say the distance from mild toxicity to recovery equals the distance from treatment failure to death. Avoid interpretations that smuggle interval-scale meaning into an ordinal endpoint.
A Reviewer’s Decision Rule
- Reconstruct the categories. Can you place every participant using prespecified, reproducible definitions and time windows?
- Challenge the ordering. Are all adjacent ranks clinically defensible, and which transitions depend on patient preferences?
- Inspect sensitivity. Does the conclusion survive reasonable alternative ranks or partial-credit schemes?
- Demand the components. Are mortality, efficacy, toxicity, and other key outcomes shown separately by group?
- Read the interval. Is the DOOR probability estimated with uncertainty, and does the interval exclude clinically unimportant differences?
- Match the design. In observational data, were treatment strategies, confounding control, positivity, and missing outcomes handled credibly? An elegant ordinal estimand does not repair a biased comparison.
Use DOOR when the clinical decision genuinely requires a joint view of benefit and harm and a defensible ordering can be prespecified. Keep outcomes separate when no coherent patient-level rank exists. If reasonable stakeholders reverse the ordering, the disagreement is part of the result and should be reported, not averaged away.
Why This Matters for Aqrab
Benefit–risk methods are a perfect test of methodological judgment. The formula can be correct while the outcome hierarchy is clinically indefensible, post hoc, or too opaque to review. Aqrab is built to trace a summary claim back to the definitions and assumptions that make it meaningful.
Paste a trial protocol or manuscript into Aqrab Try and ask whether the ranked outcome reflects a prespecified patient-relevant trade-off or merely compresses several endpoints into one favorable number. For repeatable review workflows, explore the developer tools.
Methods Anchors
The original DOOR/RADAR proposal is described by Evans and colleagues. A later randomized-trial analysis illustrates partial-credit and preference sensitivity analyses. A 2026 methodological preprint extends covariate-adjusted causal estimation of DOOR probability across randomized and observational settings; it should be read as a preprint until peer review. FDA guidance on multiple endpoints in clinical trials provides the broader multiplicity context.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Endpoint Adjudication: When a Blinded Committee Cannot Rescue Biased Event Capture
A practical endpoint adjudication guide for clinical researchers. Learn why blinded central review cannot repair unequal event capture, incomplete evidence dossiers, or post hoc endpoint rules.
Baseline Adjustment in Randomized Trials: Why Change From Baseline Keeps Losing to ANCOVA
A practical guide to baseline adjustment in randomized trials. Covers ANCOVA versus change scores, percent change traps, responder thresholds, and what reviewers should demand before trusting a tidy efficacy claim.
Win Ratio: When a Hierarchical Composite Endpoint Sounds Harder Than It Really Is
A practical guide to win ratio for clinical researchers. Covers hierarchical composite endpoints, pairwise priorities, soft-tier distortion, and what reviewers should demand before trusting a prioritized endpoint headline.