Discordant Clinical Endpoints: When One Study Produces Two Treatment Winners
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Discordant clinical endpoints are not a nuisance to smooth away. They are evidence that a study asked several questions. If relapse, disability progression, improvement, imaging, and discontinuation point in different directions, the disciplined response is not to count wins. It is to decide which outcome answers which clinical decision—and whether each estimate was measured and identified well enough to carry that claim.
A new multinational registry study comparing ofatumumab with ocrelizumab in relapsing multiple sclerosis makes the problem concrete. The target-trial emulation matched 5,288 patients. Ocrelizumab was associated with fewer relapses, while ofatumumab was associated with lower hazards of confirmed disability progression and progression independent of relapse activity. Confirmed disability improvement pointed back toward ocrelizumab; MRI activity and discontinuation did not differ significantly. That is a dashboard, not a podium.
Read Each Endpoint as Its Own Question
| Observed endpoint | Direction in the published analysis | Question it answers |
|---|---|---|
| Annualised relapse rate | Lower with ocrelizumab | How often were recorded relapses occurring? |
| Confirmed disability progression | Lower hazard with ofatumumab | How quickly did confirmed worsening occur? |
| Confirmed disability improvement | Lower hazard with ofatumumab | How quickly did confirmed improvement occur? |
| MRI activity and discontinuation | No significant difference reported | Was there evidence of separation on these domains? |
The clean metaphor
A set of endpoints is a dashboard, not a medal table. Each dial measures a different system. Adding up which side “won” more dials does not tell you which treatment is better.
Interactive endpoint-concordance audit
Turn a mixed dashboard into a defensible sentence
Use the published MS registry results as a reviewer would. Choose the decision first, then test whether the endpoint hierarchy and measurement process earn the claim.
Do not crown a winner
Evidence in view
Relapse, progression, improvement, MRI activity, and discontinuation do not all point in one direction.
Defensible sentence
The endpoint set does not identify a single overall winner. A treatment choice requires a prespecified value structure plus safety, burden, preference, and longer-term evidence.
Questions that remain
- • Who decided how relapse prevention, disability trajectories, treatment burden, safety, and patient preferences should be traded off?
- • Treat emphasis across endpoints as exploratory until the protocol or analysis plan shows a prespecified hierarchy.
- • Matching baseline covariates does not equalize how often outcomes are sought, recorded, confirmed, or missed.
Discordance Can Be Real Without Naming an Overall Winner
Relapses, progression, and improvement are related, but they are not interchangeable. They differ in event definition, time scale, baseline opportunity, confirmation rules, and clinical meaning. A treatment could plausibly reduce inflammatory relapses without producing the same short-term pattern in disability. Conversely, an apparent disability signal could emerge from measurement timing, case mix, recovery opportunity, or residual treatment channeling.
The right first move is therefore not to force biological coherence. Write one estimand sentence per endpoint: population, treatment strategies, outcome, follow-up, intercurrent-event handling, and summary measure. Only then ask whether the directions form a clinically coherent pattern.
Reviewer red flag
The abstract says one treatment had “better effectiveness” even though the endpoint hierarchy, effect scales, and clinical priorities do not support a single ordering.
Multiplicity Is Only One Failure Mode
With many endpoints, a favorable result becomes easier to find by chance. Formal multiplicity control matters most when a study makes confirmatory claims across a family of hypotheses. In observational comparative-effectiveness research, the equally important danger is selective emphasis: several outcomes are estimated, but the conclusion promotes whichever association sounds most decisive.
In this example, the abstract identifies annualised relapse rate and time to first relapse as primary outcomes; disability, MRI, and discontinuation outcomes are secondary. That does not make secondary outcomes clinically unimportant. It does mean they should extend the primary story rather than silently replace it when the directions become more interesting.
Do not audit this by counting p-values. Find the protocol or analysis plan. Identify primary, secondary, and exploratory endpoints; confirm when the hierarchy was fixed; inspect confidence intervals and effect sizes; and check whether the conclusion follows that hierarchy. A prespecified secondary endpoint can still be clinically important. It simply should not be retroactively dressed as the one question the study was always built to answer.
Matching Patients Does Not Match Measurement
Propensity-score matching can balance measured baseline covariates. It cannot ensure that two treatment groups have the same clinic schedules, MRI frequency, relapse-reporting behavior, disability-assessment completeness, or threshold for confirming an event. Nor can it remove treatment channeling driven by prognosis, access, clinician preference, prior therapy, or features absent from the registry.
This matters especially when endpoints require repeated observation. A patient must attend assessments to have worsening or improvement confirmed. If follow-up intensity or missingness differs by treatment, the opportunity to become an event can differ too. Review visit density, assessment completeness, confirmation windows, loss to follow-up, and registry-by-treatment patterns before interpreting discordance as pharmacological.
Put Relative Effects Back on an Absolute Clock
The matched study reported very low event rates and median follow-up of 1.2 years for ofatumumab and 1.4 years for ocrelizumab. The observed annualised relapse rates were 0.07 and 0.04—an absolute difference of 0.03 recorded relapses per person-year—alongside a rate ratio of 1.75. Both summaries are true descriptions of the analysis, but they create different impressions.
For every endpoint, ask for absolute risk or rate at a clinically relevant horizon, uncertainty, number at risk, and the effect of unequal follow-up. Hazard ratios for progression and improvement are not percentages of patients helped or harmed. Short follow-up can be informative about near-term disease activity while remaining insufficient for a stable overall treatment ranking.
A Five-Part Endpoint-Concordance Audit
- Name the decision. Relapse suppression, prevention of disability worsening, recovery, tolerability, and overall treatment choice are different decisions.
- Reconstruct the hierarchy. Label each endpoint primary, secondary, or exploratory from a dated protocol—not from prominence in the abstract.
- Align the estimands. Check population, treatment strategy, time horizon, outcome definition, intercurrent events, and effect measure for every endpoint.
- Audit observation opportunities. Compare visit frequency, MRI use, confirmation windows, missing assessments, and discontinuation across treatments and registries.
- Write the narrowest conclusion first. State domain-specific associations. Reserve an overall winner for evidence with an explicit, clinically justified value structure.
The safe reading of this MS study is not that one drug won. It is that the analysis found modest, directionally different associations across clinically distinct outcomes, with low event rates and limited follow-up. Those signals are useful. Their disagreement is part of the result.
Why This Matters for Aqrab
Methodology critique earns trust when it prevents a polished summary from outrunning the study beneath it. Aqrab should not merely notice that propensity matching and Cox regression were used. It should trace each treatment claim to its endpoint hierarchy, estimand, measurement process, absolute scale, and uncertainty.
Use Aqrab Try to ask whether a manuscript's conclusion respects discordant endpoints or quietly turns a mixed dashboard into a winner. For teams building repeatable appraisal workflows, the developer tools can help make that audit systematic.
Methods Anchors
The applied results come from Zhu and colleagues' multinational target-trial emulation in the Journal of Neurology, Neurosurgery & Psychiatry. A separate MS registry benchmark study shows why target-trial emulations should be validated against randomized evidence where possible. The FDA guidance on multiple endpoints explains endpoint grouping, ordering, and multiplicity in confirmatory trials; its hierarchy is useful context, not a regulatory template for declaring observational associations causal.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Informative Visit Processes: When Who Shows Up Starts Writing the Results
A practical guide to informative visit processes for clinical researchers. Covers endogenous follow-up, unequal observation schedules, visit-triggered outcome capture, inverse-intensity thinking, and what reviewers should demand before trusting longitudinal real-world results.
Delayed Entry in Survival Analysis: Why Nobody Is at Risk Before They Enter
A practical guide to left truncation in clinical survival analysis. Align the origin, entry time, event time, and risk set before interpreting Kaplan–Meier curves or Cox models from prevalent cohorts.
Target Trial Emulation Cannot Randomize Clinical Judgment: A Pertussis Study Audit
A practical target-trial emulation audit using an infant pertussis study. Check clinical-judgment confounding, propensity-score overlap, endpoint timing, sparse outcomes, and claim strength.