← Back to Blog
Evidence SynthesisGuideline MethodsMethods Critique

Transitivity in Network Meta-Analysis: When Indirect Comparisons Pretend the Trials Were Exchangeable

July 2, 2026·16 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Network meta-analysis is attractive for a simple reason: clinical decisions rarely involve only two options. Reviewers, guideline panels, and evidence syntheses often want to compare several interventions at once, including pairs that were never tested head to head.

The trouble is that many readers confuse a connected network with a trustworthy one. A shared comparator does not make indirect evidence automatically portable. Transitivity is the judgment that the linked trial sets are similar enough in the effect modifiers that matter, so an indirect comparison is not forced to travel farther than the science can support.

The Core Decision Rule

Do not ask only whether A and B each touch C somewhere in the network. Ask whether the A versus C trials and the B versus C trials are similar enough that a patient could plausibly have been randomized across the same intervention menu.

Decision rule:

If the linked comparisons differ in baseline severity, prior treatment history, outcome definition, background care, or follow-up in ways that change treatment response, the indirect estimate is not a clean substitute for a head-to-head trial.

This is why transitivity is not a line-item formality. It is the design assumption underneath every indirect comparison and every ranking built on top of it.

What Transitivity Really Means

More than a common comparator

Placebo, standard care, or “usual therapy” can connect studies structurally while hiding very different patient populations and cointerventions.

About effect modifiers

Baseline severity, prior treatment failure, dose intensity, background steroids, and follow-up timing can all change the apparent relative effect.

Prior to statistics

Node-splitting and incoherence checks matter, but they cannot rescue a network that never made clinical sense as an indirect comparison in the first place.

A Concrete Clinical Example

Case

Comparing two rheumatoid arthritis therapies through placebo when the patients are not remotely the same

Imagine a network meta-analysis comparing biologic A and oral agent B through placebo. The biologic A trials enrolled patients with severe rheumatoid arthritis after prior DMARD failure, allowed stable corticosteroids, and measured ACR50 at 24 weeks. The oral agent B trials enrolled earlier-stage, biologic-naive patients with lower inflammatory burden and measured a softer response endpoint at 12 weeks.

On a network diagram, both sets touch placebo. In a methods critique, that is the beginning of the problem, not the end. If treatment effect depends on refractoriness, baseline severity, rescue therapy, or follow-up horizon, the placebo-relative effects are being transported across a clinical divide that the diagram does not show.

The indirect A versus B estimate may still be reported. It just should not be spoken about as if the network discovered a hidden head-to-head trial.

Interactive transitivity stress test

Could these trial populations have realistically been randomized across the same menu of options?

This quick check is not a formal NMA diagnostic. It is a reviewer tool for separating a genuine indirect comparison from a network that only looks connected on paper because the comparator labels happen to match.

Current readProceed only with cautious language

How to read this network

The network may still be usable, but readers should be shown the effect-modifier imbalances explicitly and rankings should not carry the tone of settled fact.

Transitivity is a design judgment before it becomes a statistical one. If the linked trials could not plausibly have randomized the same kinds of patients among the same intervention choices, the indirect estimate is carrying more imagination than evidence.

Decision rule: treat a common comparator as a bridge only after checking the effect modifiers that decide whether response to A versus C can meaningfully inform response to B versus C.

Reviewer prompts

  • Ask whether baseline severity or risk differs enough that placebo-relative effects would change across comparisons.
  • Ask whether prior biologic exposure, line of therapy, or rescue-treatment access makes the trial sets non-exchangeable.
  • Ask whether background steroids, standard care, or dose titration differ enough to distort the indirect path.

Where Transitivity Usually Breaks

Failure modeWhat goes wrongWhat to demand instead
The comparator label matches, but the patients do notA versus placebo trials enrolled biologic-naive moderate disease, while B versus placebo trials enrolled refractory severe disease. The network looks connected, but the placebo-relative effects are being asked to travel across different clinical worlds.Show the distribution of plausible effect modifiers across each direct comparison and say plainly when indirect travel is only weakly defensible.
Different outcomes are merged into one comforting nodeProgression-free survival, symptom response, and composite remission are discussed as if they are interchangeable because all of them signal improvement.Keep outcome definitions and follow-up windows aligned, or split the network question instead of hiding the mismatch inside one pooled estimate.
Rankings are treated as more trustworthy than the network that produced themThe manuscript highlights SUCRA or rank probabilities even though transitivity concerns and incoherence diagnostics were barely addressed.Treat rankings as the last thing to trust, not the first, especially when the indirect paths are assumption-heavy.
Inconsistency testing is used as a substitute for design judgmentAuthors report no statistically significant node-splitting result and present that as proof the indirect comparison is valid, despite obvious effect-modifier imbalance and sparse loops.Use statistical inconsistency checks as a follow-up, not as a replacement for clinical examination of transitivity.

Transitivity Before Inconsistency

Authors often rush to a statistical comfort blanket here. They report that a node-splitting test or global inconsistency model did not produce an alarming p-value, then move on as if the network is now licensed for strong clinical claims.

What that cannot prove

Sparse loops, wide intervals, and few direct comparisons can make inconsistency hard to detect even when the design logic is weak.

What better reporting looks like

Show the effect modifiers first, then report agreement between direct and indirect evidence where both exist, and only then discuss rankings with appropriate restraint.

In short: no evidence of inconsistency is not the same sentence as good evidence for transitivity.

Why This Matters for Guidelines and HTA

Network meta-analysis is especially tempting in guideline panels and health technology assessment because it promises a full comparative picture, not just isolated pairwise fragments. That promise is worth pursuing. It also creates a specific failure mode: certainty gets borrowed from the elegance of the network diagram rather than from the comparability of the underlying trials.

If the transitivity story is weak, treatment rankings and indirect superiority claims should be written with more caution than the graphic often suggests. This is exactly the kind of methodological pressure point Aqrab can help surface during review, especially when a paper's clean visual summary is doing more persuasive work than its design logic.

Reviewer Red-Flag Checklist

  • Could the patients in each linked comparison plausibly have been randomized across the same intervention menu?
  • Which trial-level effect modifiers matter here: baseline severity, prior lines of therapy, background care, dose strategy, follow-up horizon, or endpoint definition?
  • Where both direct and indirect evidence exist, did the authors check agreement rather than skipping straight to the network estimate?
  • Are treatment rankings being emphasized more strongly than the assumptions needed to generate them?
  • Would the clinical conclusion still sound persuasive if the indirect evidence were described without the ranking graphic?

The Practical Bottom Line

A network meta-analysis becomes useful when the indirect paths are clinically believable, the direct and indirect evidence are checked against each other, and the resulting claims are scaled to the strength of those assumptions. Without that discipline, the network is often answering a cleaner question than the included trials were ever designed to support.

If you want a fast way to stress-test whether an indirect comparison is behaving like evidence or like architecture, Aqrab can help surface effect-modifier mismatches, ranking overclaim, and weak comparator logic before the manuscript leaves your desk. If you want those checks inside your own review workflow, the developer tools are the more scalable place to start.

References and Further Reading

  • Chaimani A, Caldwell DM, Li T, Higgins JPT, Salanti G. Undertaking network meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5; 2024 update.
  • Jansen JP, Naci H. Is network meta-analysis as valid as standard pairwise meta-analysis? It all depends on the distribution of effect modifiers. BMC Medicine. 2013;11:159.
  • Puhan MA, Schunemann HJ, Murad MH, et al. A GRADE Working Group approach for rating the quality of treatment effect estimates from network meta-analysis. BMJ. 2014;349:g5630.
  • Hutton B, Salanti G, Caldwell DM, et al. The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions. Ann Intern Med. 2015;162:777-784.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive