Win Ratio: When a Hierarchical Composite Endpoint Sounds Harder Than It Really Is
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Some trials do not want to treat death, hospitalization, and symptom change as interchangeable pieces of a flat composite. That instinct is sensible. Patients do not rank those outcomes equally, so analysts reach for the win ratio: compare treated and control patients in pairs, judge the hardest outcome first, and move to lower tiers only when the higher tier ties.
In the best cases, that hierarchy makes the estimand more clinically honest than an ordinary composite. In the worst cases, it produces a headline that sounds like a mortality result even though most wins came from softer tie-breakers far downstream in the ranking.
The Core Decision Rule
Do not ask only whether the hierarchy is clinically plausible. Ask whether the positive signal would still feel persuasive if you reported where the wins actually came from.
Decision rule:
A win ratio earns trust when the priority order is prespecified, the upper tiers matter to patients, and the treatment advantage is not mostly manufactured after the analysis falls through repeated hard outcome ties.
The crucial distinction is between respecting patient priorities and borrowing prestige from them. Those are not the same thing.
Why Investigators Reach for Win Ratio
It avoids flat composites
Death can outrank hospitalization, and hospitalization can outrank symptom drift, instead of all components pretending to carry equal clinical weight.
It can use more information
When deaths are uncommon, the method can still compare patients on the next clinically meaningful tier instead of declaring most pairs permanently tied.
It sounds patient-centered
That part is true only if the ranking, definitions, and reporting stay disciplined after the abstract is written.
A Concrete Clinical Example
Case
Advanced heart failure trial with death, hospitalization, then symptom score
Imagine a trial where treated and control patients are paired on baseline severity. Each pair is compared first on time to death. If neither patient clearly wins there, the comparison falls to heart failure hospitalization. If that also ties, the last tie-breaker is a symptom-status score.
That setup can be sensible. It can also become rhetorically slippery. If death is nearly unchanged, hospitalization moves modestly, and most wins come from symptom score, the trial may still report a favorable win ratio that sounds tougher than the hardest observed outcomes really were.
The paper is not necessarily wrong. It just owes you more than one ratio and a triumphant abstract sentence.
Interactive win-ratio explorer
See when the hierarchy is powered by hard outcomes versus soft tie-breakers
This is a teaching device, not a replacement for a prespecified win-ratio analysis. It turns the method into 1,000 hypothetical treatment-versus-control pairwise comparisons and asks which tier actually creates the positive headline.
Higher values mean more pairwise contests are decided by the hardest endpoint instead of lower-tier ties.
This approximates how often the analysis falls to a softer but still clinically serious tie-breaker.
When this climbs, the headline depends more heavily on lower-priority patient-status differences.
Treatment wins
424 pairs
Control wins
257 pairs
Wins from death tier
15%
Wins from lower tiers
85%
How to read this setup
The result does not look absurd on its face, but you still need the component-level story before treating the headline as clinically decisive.
A win ratio can be genuinely useful when the hierarchy mirrors clear patient priorities and the upper tiers still carry real information. It becomes much less persuasive when the hardest outcome mostly ties and the positive result is manufactured downstream.
Tier-by-tier accounting
- Death tier: 120 pairs decisive pairs, with 64 pairs wins for treatment.
- Hospitalization tier: 299 pairs decisive pairs, with 183 pairs wins for treatment.
- Symptom tier: 261 pairs decisive pairs, with 178 pairs wins for treatment.
- Still tied after all tiers: 319 pairs.
| Pattern | What it implies | Reviewer reaction |
|---|---|---|
| Positive ratio, but death tier near neutral | The hierarchy may be borrowing gravitas from the top tier while getting its signal elsewhere. | Demand component-level counts and explain the clinical weight of the downstream wins. |
| Positive ratio with meaningful upper-tier contribution | The result is more aligned with the stated patient-priority hierarchy. | Still check definitions, adjudication, and whether the ranking was prespecified. |
| Most pairs end unresolved | The method may be data-hungry for this endpoint stack and vulnerable to unstable interpretation. | Ask whether a simpler primary endpoint or a time-to-event estimand would have been cleaner. |
Where Win Ratio Starts to Mislead
| Failure mode | What goes wrong | Why reviewers should care |
|---|---|---|
| Upper tiers mostly tie | The result is carried by downstream tiers while still wearing the prestige of death-first prioritization. | A positive ratio may sound more clinically severe than the actual signal source. |
| Lower tiers are subjective or practice-sensitive | Symptom scales, urgent visits, or treatment intensification can be influenced by ascertainment and site culture. | The hierarchy becomes vulnerable to open-label behavior and measurement drift. |
| Pairing logic is opaque | Readers cannot see how patients were matched or how ties were adjudicated. | Interpretation becomes inseparable from implementation details the manuscript did not earn. |
| The hierarchy changed after seeing data | Analysts can quietly move clinically weaker outcomes upward or add extra tie-breakers. | That is endpoint gardening with better manners. |
When the Method Is Actually a Good Fit
More defensible
- The hierarchy mirrors obvious patient priorities.
- Upper tiers still decide a meaningful share of pairwise comparisons.
- Definitions are objective or tightly adjudicated.
- The protocol prespecified ranking, pairing, and tie rules.
Less defensible
- The top tier rarely contributes and mostly serves as branding.
- Lower tiers are subjective, site-dependent, or operationally noisy.
- Component reporting is sparse because the ratio alone looks cleaner.
- The manuscript implies the hierarchy solved every problem a flat composite had.
Reviewer Red-Flag Checklist
- Ask what proportion of pairwise wins was decided by each tier, not only the overall ratio.
- Ask whether the hierarchy was prespecified in the protocol and SAP before unblinding.
- Ask whether subjective lower-tier outcomes were assessed equally across groups.
- Ask whether a simpler primary endpoint would have been more clinically transparent.
- Ask whether the discussion section speaks as if death improved when the data mostly say symptoms improved.
What to Report Instead of a Lone Headline Ratio
A serious win-ratio paper should show the hierarchy, the matching logic, the proportion of pairs resolved at each tier, and component-level summaries that let readers see what clinical question was actually answered. Without that accounting, the ratio is too compressed for the rhetorical work it is being asked to do.
This is exactly the sort of design choice that becomes stronger under explicit methodological review. If you want a live protocol, SAP, or manuscript endpoint strategy pressure-tested before submission, Aqrab can review the hierarchy, estimand, and reporting logic against the question you are really trying to answer. Start with the Aqrab trial critique workflow when the endpoint story feels elegant on paper but harder to defend in review.
The Bottom Line
Win ratio is not methodological theater by default. It can be a thoughtful way to respect clinical priorities. But a prioritized endpoint is only as honest as the tiers that actually carry the result.
If the ratio sounds like a hard-outcome triumph while the real leverage came from lower-tier tie-breakers, the right response is not admiration for analytical sophistication. It is a request for a more honest clinical summary.