← Back to Blog
Survival AnalysisClinical TrialsMethods Critique

Nonproportional Hazards: When One Hazard Ratio Pretends the Treatment Effect Never Changes

July 5, 2026·16 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Survival-analysis papers often report a hazard ratio as though the treatment effect stayed the same from the first event to the last. Real trials are rarely that polite. Toxicity can hurt early and help later. Immunotherapy can look idle before curves finally separate. Crossover, waning adherence, or rescue therapy can flatten an initially strong signal.

The practical problem is not merely that the proportional hazards assumption failed on a technical exam. It is that one pooled hazard ratio can stop matching any clinically recognizable period of the trial. Once that happens, the question is no longer whether the number is estimable. It is whether the number is still an honest summary.

The Core Decision Rule

Treat a single hazard ratio as a thin summary when the Kaplan-Meier curves cross, separate only after a delay, or reconverge later in follow-up. In those settings, the paper should show the curve, state what the clinically important time window is, and add a complementary summary such as RMST or a pre-specified landmark rate.

Decision rule:

If the treatment story changes over time, do not let the abstract imply that one hazard ratio is the whole treatment story.

A hazard ratio is most comfortable when the relative effect is reasonably stable. When the clinical tradeoff evolves, the analyst needs to show the evolution rather than compress it into one confident line.

What Nonproportional Hazards Actually Mean

The hazard ratio is no longer stable over time

Under nonproportional hazards, the relative event rate can change meaningfully from early to late follow-up. A pooled Cox estimate then becomes a weighted average of changing periods rather than one clean clinical contrast.

The problem is interpretability, not only fit

A manuscript can pass routine model diagnostics and still communicate badly if the early and late parts of the trial imply different decisions.

The usual log-rank default can lose sharpness

Standard time-to-event tests are most comfortable when hazards stay proportional. Once curves cross or effects are delayed, power and interpretability can deteriorate together.

The protocol should anticipate this possibility

If delayed effects or crossing curves are plausible, supplementary analyses and follow-up plans should be prespecified instead of improvised after the graph becomes awkward.

A Concrete Clinical Example

Case

Adjuvant oncology trial with early toxicity and later durable disease control

Imagine a randomized trial of an adjuvant immune therapy. During the first few months, immune-related toxicities and discontinuations make event-free survival look worse than control. By month 12, the treated curve starts to climb above control because recurrences are being delayed in the patients who tolerated and benefited from treatment.

A single hazard ratio can report a statistically respectable average across those eras, but clinicians still need to know three separate things: how much early harm occurred, when the curve crossed, and whether the later benefit was large enough to justify the early tradeoff.

That is the real reading task. The hazard ratio is only part of it.

Interactive survival-analysis explorer

When does one hazard ratio stop being an honest summary?

This teaching tool shows why the same trial can look early harmful, late beneficial, or merely ambiguous depending on follow-up maturity and whether alternative summaries were planned before the curves appeared.

Single-HR credibility31/100One pooled ratio is likely hiding too much.

Toxicity, learning curves, or implementation friction hurt early before later benefit appears.

Low maturity means the paper is speaking before the late part of the curve has had much chance to matter.

Higher values mean the protocol anticipated non-proportional hazards instead of improvising after the curves crossed.

Log-rank reliability

36/100

How comfortable you should be with a default time-to-event comparison.

RMST usefulness

80/100

Estimated event-free time difference at the current horizon: +0.7 months.

Milestone summary value

75/100

Useful when a fixed clinical horizon matters more than a pooled average over the entire follow-up.

Curve story

Early:

Early months favor control because the intervention carries upfront cost or risk.

Late:

Later months favor treatment once the durable benefit arrives.

What to report

Report the Kaplan-Meier curve, a prespecified RMST horizon, and a plain-language account of when the tradeoff flips.

Reviewer checkpoint

Ask whether the abstract is hiding a clinically important early-harm period behind a favorable average.

If the pattern looks like thisDo not stop atAdd this
A single headline hazard ratio compresses two different clinical eras into one number.A lone hazard ratio in the abstractKaplan-Meier curves, a pre-specified complementary summary, and time-specific clinical interpretation
Protocol did not anticipate non-PHPost hoc metric-shopping after the curves became awkwardTransparent labeling of exploratory analyses and a conservative claims section
Clinicians care about a fixed decision windowGeneric “better survival” languageA milestone rate or RMST difference at that clinical horizon

Four Patterns That Create Nonproportional Hazards

1. Delayed benefit

Immunologic, behavioral, and implementation effects may need time before curves separate. Early follow-up can therefore understate the eventual benefit.

2. Early harm with later benefit

Toxicity, procedural risk, and learning curves can produce a clinically important early penalty that should not vanish inside a later average.

3. Waning effect

Early benefits can shrink when adherence fades, crossover accumulates, or the intervention acts only on short-term mechanisms.

4. Hidden treatment-version or post-randomization change

Subsequent therapy, rescue protocols, and evolving standard care can alter hazards over time even if the original randomization was clean.

What to Report Instead of Acting as Though the Problem Does Not Exist

QuestionHelpful summaryWhy it helps
What did patients gain over a fixed window?Restricted mean survival timeTurns evolving survival curves into time gained or lost by a clinically meaningful horizon.
What is the status at a prespecified clinical milestone?Landmark survival or event-free rateUseful when a specific decision window matters more than an overall average.
Did the effect change over time?Kaplan-Meier curve plus time-aware interpretationShows whether the tradeoff is early, late, transient, or persistent.
Was this planned prospectively?Pre-specified sensitivity or supplementary analysesProtects the paper from looking like it changed summaries only after the hazard ratio became inconvenient.

Five Reviewer Red Flags

1. The curves cross, but the abstract sounds monotone

If early harm and later benefit are both present, a one-direction conclusion is usually telling only half the clinical story.

2. The paper reports one hazard ratio and no compensating summary

Readers need a time-based quantity or milestone estimate once the effect clearly evolves over follow-up.

3. Alternative analyses appear only after the figure became awkward

Post hoc metric-shopping is still metric-shopping even when the new metric is more sensible.

4. Follow-up is too immature for the claimed long-term story

Delayed benefit and late toxicity cannot be read confidently from a trial that stopped talking too early.

5. Intercurrent events are shaping the curve, but the paper never says how

Crossover, subsequent therapy, and treatment discontinuation can all change hazard patterns over time and deserve explicit interpretation.

What Reviewers Should Demand Before Trusting the Headline

  • The Kaplan-Meier curve with numbers at risk rather than only a hazard ratio and p-value.
  • A statement about whether delayed effects or curve crossing were plausible at the design stage.
  • Pre-specified checks or supplementary analyses for nonproportional hazards.
  • A clinically interpretable time-based summary when the effect changes over follow-up.
  • Plain-language interpretation of early harm, late benefit, or waning benefit instead of one-direction spin.

Where Aqrab Fits

Nonproportional hazards are exactly the kind of methods problem that can hide inside polished survival language. Aqrab is useful when you want the manuscript pressure-tested for whether the curve, the summary measure, the estimand, and the conclusion are actually aligned.

If you are reviewing a survival paper with crossing curves or delayed effects, the Aqrab critique flow is the fastest way to see whether the headline claim is sturdier than the hazard ratio alone.

The Bottom Line

Nonproportional hazards are not a nuisance footnote. They are a warning that the treatment effect may have different meanings at different times in the same trial.

Once that happens, the right response is not to worship a single ratio more forcefully. It is to report the time pattern honestly and choose summaries that still sound like the clinical question you meant to answer.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive