Transitivity in Network Meta-Analysis: When Indirect Comparisons Pretend the Trials Were Exchangeable
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Network meta-analysis is attractive for a simple reason: clinical decisions rarely involve only two options. Reviewers, guideline panels, and evidence syntheses often want to compare several interventions at once, including pairs that were never tested head to head.
The trouble is that many readers confuse a connected network with a trustworthy one. A shared comparator does not make indirect evidence automatically portable. Transitivity is the judgment that the linked trial sets are similar enough in the effect modifiers that matter, so an indirect comparison is not forced to travel farther than the science can support.
The Core Decision Rule
Do not ask only whether A and B each touch C somewhere in the network. Ask whether the A versus C trials and the B versus C trials are similar enough that a patient could plausibly have been randomized across the same intervention menu.
Decision rule:
If the linked comparisons differ in baseline severity, prior treatment history, outcome definition, background care, or follow-up in ways that change treatment response, the indirect estimate is not a clean substitute for a head-to-head trial.
This is why transitivity is not a line-item formality. It is the design assumption underneath every indirect comparison and every ranking built on top of it.
What Transitivity Really Means
More than a common comparator
Placebo, standard care, or “usual therapy” can connect studies structurally while hiding very different patient populations and cointerventions.
About effect modifiers
Baseline severity, prior treatment failure, dose intensity, background steroids, and follow-up timing can all change the apparent relative effect.
Prior to statistics
Node-splitting and incoherence checks matter, but they cannot rescue a network that never made clinical sense as an indirect comparison in the first place.
A Concrete Clinical Example
Case
Comparing two rheumatoid arthritis therapies through placebo when the patients are not remotely the same
Imagine a network meta-analysis comparing biologic A and oral agent B through placebo. The biologic A trials enrolled patients with severe rheumatoid arthritis after prior DMARD failure, allowed stable corticosteroids, and measured ACR50 at 24 weeks. The oral agent B trials enrolled earlier-stage, biologic-naive patients with lower inflammatory burden and measured a softer response endpoint at 12 weeks.
On a network diagram, both sets touch placebo. In a methods critique, that is the beginning of the problem, not the end. If treatment effect depends on refractoriness, baseline severity, rescue therapy, or follow-up horizon, the placebo-relative effects are being transported across a clinical divide that the diagram does not show.
The indirect A versus B estimate may still be reported. It just should not be spoken about as if the network discovered a hidden head-to-head trial.
Interactive transitivity stress test
Could these trial populations have realistically been randomized across the same menu of options?
This quick check is not a formal NMA diagnostic. It is a reviewer tool for separating a genuine indirect comparison from a network that only looks connected on paper because the comparator labels happen to match.
How to read this network
The network may still be usable, but readers should be shown the effect-modifier imbalances explicitly and rankings should not carry the tone of settled fact.
Transitivity is a design judgment before it becomes a statistical one. If the linked trials could not plausibly have randomized the same kinds of patients among the same intervention choices, the indirect estimate is carrying more imagination than evidence.
Decision rule: treat a common comparator as a bridge only after checking the effect modifiers that decide whether response to A versus C can meaningfully inform response to B versus C.
Reviewer prompts
- Ask whether baseline severity or risk differs enough that placebo-relative effects would change across comparisons.
- Ask whether prior biologic exposure, line of therapy, or rescue-treatment access makes the trial sets non-exchangeable.
- Ask whether background steroids, standard care, or dose titration differ enough to distort the indirect path.
Where Transitivity Usually Breaks
| Failure mode | What goes wrong | What to demand instead |
|---|---|---|
| The comparator label matches, but the patients do not | A versus placebo trials enrolled biologic-naive moderate disease, while B versus placebo trials enrolled refractory severe disease. The network looks connected, but the placebo-relative effects are being asked to travel across different clinical worlds. | Show the distribution of plausible effect modifiers across each direct comparison and say plainly when indirect travel is only weakly defensible. |
| Different outcomes are merged into one comforting node | Progression-free survival, symptom response, and composite remission are discussed as if they are interchangeable because all of them signal improvement. | Keep outcome definitions and follow-up windows aligned, or split the network question instead of hiding the mismatch inside one pooled estimate. |
| Rankings are treated as more trustworthy than the network that produced them | The manuscript highlights SUCRA or rank probabilities even though transitivity concerns and incoherence diagnostics were barely addressed. | Treat rankings as the last thing to trust, not the first, especially when the indirect paths are assumption-heavy. |
| Inconsistency testing is used as a substitute for design judgment | Authors report no statistically significant node-splitting result and present that as proof the indirect comparison is valid, despite obvious effect-modifier imbalance and sparse loops. | Use statistical inconsistency checks as a follow-up, not as a replacement for clinical examination of transitivity. |
Transitivity Before Inconsistency
Authors often rush to a statistical comfort blanket here. They report that a node-splitting test or global inconsistency model did not produce an alarming p-value, then move on as if the network is now licensed for strong clinical claims.
What that cannot prove
Sparse loops, wide intervals, and few direct comparisons can make inconsistency hard to detect even when the design logic is weak.
What better reporting looks like
Show the effect modifiers first, then report agreement between direct and indirect evidence where both exist, and only then discuss rankings with appropriate restraint.
In short: no evidence of inconsistency is not the same sentence as good evidence for transitivity.
Why This Matters for Guidelines and HTA
Network meta-analysis is especially tempting in guideline panels and health technology assessment because it promises a full comparative picture, not just isolated pairwise fragments. That promise is worth pursuing. It also creates a specific failure mode: certainty gets borrowed from the elegance of the network diagram rather than from the comparability of the underlying trials.
If the transitivity story is weak, treatment rankings and indirect superiority claims should be written with more caution than the graphic often suggests. This is exactly the kind of methodological pressure point Aqrab can help surface during review, especially when a paper's clean visual summary is doing more persuasive work than its design logic.
Reviewer Red-Flag Checklist
- •Could the patients in each linked comparison plausibly have been randomized across the same intervention menu?
- •Which trial-level effect modifiers matter here: baseline severity, prior lines of therapy, background care, dose strategy, follow-up horizon, or endpoint definition?
- •Where both direct and indirect evidence exist, did the authors check agreement rather than skipping straight to the network estimate?
- •Are treatment rankings being emphasized more strongly than the assumptions needed to generate them?
- •Would the clinical conclusion still sound persuasive if the indirect evidence were described without the ranking graphic?
The Practical Bottom Line
A network meta-analysis becomes useful when the indirect paths are clinically believable, the direct and indirect evidence are checked against each other, and the resulting claims are scaled to the strength of those assumptions. Without that discipline, the network is often answering a cleaner question than the included trials were ever designed to support.
If you want a fast way to stress-test whether an indirect comparison is behaving like evidence or like architecture, Aqrab can help surface effect-modifier mismatches, ranking overclaim, and weak comparator logic before the manuscript leaves your desk. If you want those checks inside your own review workflow, the developer tools are the more scalable place to start.
References and Further Reading
- Chaimani A, Caldwell DM, Li T, Higgins JPT, Salanti G. Undertaking network meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions. Version 6.5; 2024 update.
- Jansen JP, Naci H. Is network meta-analysis as valid as standard pairwise meta-analysis? It all depends on the distribution of effect modifiers. BMC Medicine. 2013;11:159.
- Puhan MA, Schunemann HJ, Murad MH, et al. A GRADE Working Group approach for rating the quality of treatment effect estimates from network meta-analysis. BMJ. 2014;349:g5630.
- Hutton B, Salanti G, Caldwell DM, et al. The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions. Ann Intern Med. 2015;162:777-784.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Indirectness in Clinical Evidence: When a Good Study Answers the Wrong Question
A practical guide to indirectness in clinical evidence for clinical researchers. Covers PICO mismatch, outdated comparators, surrogate outcomes, and what reviewers should demand before trusting an applicable-sounding conclusion.
ROBINS-I: When an Observational Effect Estimate Is Too Biased to Grade Casually
A practical guide to ROBINS-I for clinical researchers. Covers the seven bias domains, why confounding and time zero usually dominate, how overall judgments are formed, and what reviewers should demand before trusting non-randomized evidence.
Guideline Recommendation Strength: When “Strongly Recommend” Starts Outrunning the Evidence
A practical guide to recommendation strength in clinical guidelines. Covers certainty of evidence, benefit-harm tradeoffs, patient values, implementation burden, and what reviewers should demand before trusting a forceful recommendation.