← Back to Blog
Real-World EvidenceOutcome MeasurementMethods Critique

Discordant Clinical Endpoints: When One Study Produces Two Treatment Winners

August 16, 2026·14 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Discordant clinical endpoints are not a nuisance to smooth away. They are evidence that a study asked several questions. If relapse, disability progression, improvement, imaging, and discontinuation point in different directions, the disciplined response is not to count wins. It is to decide which outcome answers which clinical decision—and whether each estimate was measured and identified well enough to carry that claim.

A new multinational registry study comparing ofatumumab with ocrelizumab in relapsing multiple sclerosis makes the problem concrete. The target-trial emulation matched 5,288 patients. Ocrelizumab was associated with fewer relapses, while ofatumumab was associated with lower hazards of confirmed disability progression and progression independent of relapse activity. Confirmed disability improvement pointed back toward ocrelizumab; MRI activity and discontinuation did not differ significantly. That is a dashboard, not a podium.

Read Each Endpoint as Its Own Question

Observed endpointDirection in the published analysisQuestion it answers
Annualised relapse rateLower with ocrelizumabHow often were recorded relapses occurring?
Confirmed disability progressionLower hazard with ofatumumabHow quickly did confirmed worsening occur?
Confirmed disability improvementLower hazard with ofatumumabHow quickly did confirmed improvement occur?
MRI activity and discontinuationNo significant difference reportedWas there evidence of separation on these domains?

The clean metaphor

A set of endpoints is a dashboard, not a medal table. Each dial measures a different system. Adding up which side “won” more dials does not tell you which treatment is better.

Interactive endpoint-concordance audit

Turn a mixed dashboard into a defensible sentence

Teaching tool · not clinical advice

Use the published MS registry results as a reviewer would. Choose the decision first, then test whether the endpoint hierarchy and measurement process earn the claim.

1. What decision are you trying to inform?
2. Was the endpoint hierarchy prespecified?
3. Is outcome ascertainment comparable?

Do not crown a winner

Evidence in view

Relapse, progression, improvement, MRI activity, and discontinuation do not all point in one direction.

Defensible sentence

The endpoint set does not identify a single overall winner. A treatment choice requires a prespecified value structure plus safety, burden, preference, and longer-term evidence.

Questions that remain

  • Who decided how relapse prevention, disability trajectories, treatment burden, safety, and patient preferences should be traded off?
  • Treat emphasis across endpoints as exploratory until the protocol or analysis plan shows a prespecified hierarchy.
  • Matching baseline covariates does not equalize how often outcomes are sought, recorded, confirmed, or missed.

Discordance Can Be Real Without Naming an Overall Winner

Relapses, progression, and improvement are related, but they are not interchangeable. They differ in event definition, time scale, baseline opportunity, confirmation rules, and clinical meaning. A treatment could plausibly reduce inflammatory relapses without producing the same short-term pattern in disability. Conversely, an apparent disability signal could emerge from measurement timing, case mix, recovery opportunity, or residual treatment channeling.

The right first move is therefore not to force biological coherence. Write one estimand sentence per endpoint: population, treatment strategies, outcome, follow-up, intercurrent-event handling, and summary measure. Only then ask whether the directions form a clinically coherent pattern.

Reviewer red flag

The abstract says one treatment had “better effectiveness” even though the endpoint hierarchy, effect scales, and clinical priorities do not support a single ordering.

Multiplicity Is Only One Failure Mode

With many endpoints, a favorable result becomes easier to find by chance. Formal multiplicity control matters most when a study makes confirmatory claims across a family of hypotheses. In observational comparative-effectiveness research, the equally important danger is selective emphasis: several outcomes are estimated, but the conclusion promotes whichever association sounds most decisive.

In this example, the abstract identifies annualised relapse rate and time to first relapse as primary outcomes; disability, MRI, and discontinuation outcomes are secondary. That does not make secondary outcomes clinically unimportant. It does mean they should extend the primary story rather than silently replace it when the directions become more interesting.

Do not audit this by counting p-values. Find the protocol or analysis plan. Identify primary, secondary, and exploratory endpoints; confirm when the hierarchy was fixed; inspect confidence intervals and effect sizes; and check whether the conclusion follows that hierarchy. A prespecified secondary endpoint can still be clinically important. It simply should not be retroactively dressed as the one question the study was always built to answer.

Matching Patients Does Not Match Measurement

Propensity-score matching can balance measured baseline covariates. It cannot ensure that two treatment groups have the same clinic schedules, MRI frequency, relapse-reporting behavior, disability-assessment completeness, or threshold for confirming an event. Nor can it remove treatment channeling driven by prognosis, access, clinician preference, prior therapy, or features absent from the registry.

This matters especially when endpoints require repeated observation. A patient must attend assessments to have worsening or improvement confirmed. If follow-up intensity or missingness differs by treatment, the opportunity to become an event can differ too. Review visit density, assessment completeness, confirmation windows, loss to follow-up, and registry-by-treatment patterns before interpreting discordance as pharmacological.

Put Relative Effects Back on an Absolute Clock

The matched study reported very low event rates and median follow-up of 1.2 years for ofatumumab and 1.4 years for ocrelizumab. The observed annualised relapse rates were 0.07 and 0.04—an absolute difference of 0.03 recorded relapses per person-year—alongside a rate ratio of 1.75. Both summaries are true descriptions of the analysis, but they create different impressions.

For every endpoint, ask for absolute risk or rate at a clinically relevant horizon, uncertainty, number at risk, and the effect of unequal follow-up. Hazard ratios for progression and improvement are not percentages of patients helped or harmed. Short follow-up can be informative about near-term disease activity while remaining insufficient for a stable overall treatment ranking.

A Five-Part Endpoint-Concordance Audit

  1. Name the decision. Relapse suppression, prevention of disability worsening, recovery, tolerability, and overall treatment choice are different decisions.
  2. Reconstruct the hierarchy. Label each endpoint primary, secondary, or exploratory from a dated protocol—not from prominence in the abstract.
  3. Align the estimands. Check population, treatment strategy, time horizon, outcome definition, intercurrent events, and effect measure for every endpoint.
  4. Audit observation opportunities. Compare visit frequency, MRI use, confirmation windows, missing assessments, and discontinuation across treatments and registries.
  5. Write the narrowest conclusion first. State domain-specific associations. Reserve an overall winner for evidence with an explicit, clinically justified value structure.

The safe reading of this MS study is not that one drug won. It is that the analysis found modest, directionally different associations across clinically distinct outcomes, with low event rates and limited follow-up. Those signals are useful. Their disagreement is part of the result.

Why This Matters for Aqrab

Methodology critique earns trust when it prevents a polished summary from outrunning the study beneath it. Aqrab should not merely notice that propensity matching and Cox regression were used. It should trace each treatment claim to its endpoint hierarchy, estimand, measurement process, absolute scale, and uncertainty.

Use Aqrab Try to ask whether a manuscript's conclusion respects discordant endpoints or quietly turns a mixed dashboard into a winner. For teams building repeatable appraisal workflows, the developer tools can help make that audit systematic.

Methods Anchors

The applied results come from Zhu and colleagues' multinational target-trial emulation in the Journal of Neurology, Neurosurgery & Psychiatry. A separate MS registry benchmark study shows why target-trial emulations should be validated against randomized evidence where possible. The FDA guidance on multiple endpoints explains endpoint grouping, ordering, and multiplicity in confirmatory trials; its hierarchy is useful context, not a regulatory template for declaring observational associations causal.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive