Indirectness in Clinical Evidence: When a Good Study Answers the Wrong Question
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Indirectness is one of the most polite ways evidence can mislead. The trial can be randomized, the estimates can be precise, and the methods section can look immaculate. But if the paper studied the wrong population, a softer intervention, an obsolete comparator, or a surrogate endpoint that stands in for the real decision, the evidence is no longer fully about your question.
That is why indirectness is not a fussy guideline footnote. It is the discipline of checking whether the evidence and the decision are still attached to each other.
The Core Decision Rule
Do not ask only whether the study was internally valid. Ask whether the study still matches the decision-maker's population, intervention, comparator, outcome, and implementation setting.
Decision rule:
If the paper would force you to add phrases like “in younger patients,” “against older usual care,” “for a surrogate only,” or “inside a much more resourced program,” you are already paying an indirectness tax and should say so explicitly.
Why Indirectness Gets Missed
Methodological polish steals attention
Readers often stop at randomization, low loss to follow-up, or model quality and never finish the harder question of applicability.
PICO drift happens one domain at a time
A paper can look nearly relevant until you notice the healthier sample, extra coaching, softer comparator, or shorter follow-up that quietly changed the question.
People treat “same disease” as enough
Sharing a disease label does not guarantee the same baseline risk, treatment tolerance, workflow, or patient-important outcomes.
A Concrete Clinical Example
Case
A digital blood-pressure program that worked inside a heavily supported academic network
Imagine a trial of remote hypertension management with pharmacist titration, automated reminders, loaned devices, and weekly outreach. The participants are relatively motivated, English-speaking, and connected to a tertiary health system. The primary endpoint is systolic blood pressure at 12 weeks.
Now imagine someone citing that paper to justify a broad implementation claim for a low-touch mobile app in under-resourced primary-care clinics, for older multimorbid patients, with no pharmacist support and no home device program.
The original study may still be good science. But the recommendation has drifted across population, intervention, setting, and possibly outcome horizon. What looked like “evidence-based rollout” is now evidence that must be narrowed, qualified, or downgraded before it can travel honestly.
Interactive indirectness stress test
A precise estimate can still be answering someone else's question
This teaching tool turns five common applicability mismatches into one visible judgment. It is not a formal GRADE worksheet. It is a way to stop “good evidence” from becoming shorthand for “direct evidence.”
The trial excluded the frailer or more complex patients you now want to treat.
The clinical idea is similar, but the package, adherence support, or implementation bundle differs.
The control arm resembles an older era or a setting where the standard pathway is weaker.
The signal is relevant, but it stands one step away from the patient outcome or time horizon you care about.
The result may still help, but it depends on support systems the target environment may not have.
Serious mismatches
0
Two or more major mismatches usually mean the paper should not travel into a strong applied claim unchanged.
Interpretation
Moderate indirectness pressure
The evidence is starting to answer a nearby question rather than the exact one. You should explain the mismatch instead of treating applicability as automatic.
| Domain | Current judgment | Why it matters |
|---|---|---|
| Population | Population is narrower or healthier | The trial excluded the frailer or more complex patients you now want to treat. |
| Intervention | Intervention is related but modified | The clinical idea is similar, but the package, adherence support, or implementation bundle differs. |
| Comparator | Comparator is dated or only partly relevant | The control arm resembles an older era or a setting where the standard pathway is weaker. |
| Outcome and horizon | Outcome is a surrogate or follow-up is shorter | The signal is relevant, but it stands one step away from the patient outcome or time horizon you care about. |
| Setting and workflow | Setting is more resourced than the target setting | The result may still help, but it depends on support systems the target environment may not have. |
This is a teaching illustration, not a replacement for full evidence appraisal. Its job is to make applicability drift visible before it gets hidden inside a confident conclusion.
The Five Indirectness Questions That Matter Most
| Domain | What to ask | Typical failure mode |
|---|---|---|
| Population | Are the patients in the paper close to the patients for whom the decision is being made? | Younger, cleaner, lower-risk participants are used to support claims in frailer real-world populations. |
| Intervention | Is the tested package the same as the intervention people plan to implement? | A complex support bundle is reduced to “the app worked” or “the drug worked” after stripping away the rest. |
| Comparator | Does the control arm still represent the real alternative clinicians face today? | A benefit against outdated usual care gets retold as a benefit against modern optimized care. |
| Outcome and horizon | Was the measured endpoint truly patient-important over a meaningful time window? | A short-term biomarker change gets promoted into a long-term clinical benefit claim. |
| Setting and workflow | Could the result survive outside the original staffing, monitoring, and adherence environment? | Specialist-center performance is assumed to generalize into lightly supported routine practice. |
When Indirectness Is Mild Versus Serious
Mild indirectness
One domain is stretched but the clinical story still mostly holds. A guideline or review can often use the study, but should narrow the claim and acknowledge why confidence is not perfect.
Moderate indirectness
The evidence remains informative, but only after translation. This is where decision-makers must show their work rather than borrowing the trial conclusion as-is.
Serious indirectness
Several domains have drifted or one domain is extreme. The evidence may still generate hypotheses, but it should not carry a strong recommendation without visible qualification.
Reviewer Red Flags Before Trusting the Applicability Claim
1. The conclusion quietly changes the comparator
If the paper beat placebo, minimal care, or an older pathway, the discussion should not speak as though it beat today's best alternative.
2. The outcome being promoted is not the one that matters clinically
Biomarkers and short follow-up windows can be useful, but they should not be smuggled into claims about long-term patient benefit without argument.
3. Support infrastructure disappears in the retelling
Coaching, monitoring, adjudication, pharmacist support, and specialist follow-up are often part of the intervention, not background scenery.
4. The paper's exclusions are treated like trivial housekeeping
If the trial excluded the very patients most likely to receive the intervention in practice, the applicability claim should start narrow, not broad.
What Better Evidence Writing Looks Like
Better writing does not pretend indirectness away. It says, for example, that the evidence supports a pharmacist-supported hypertension program in relatively engaged patients over 12 weeks, not that any digital blood-pressure intervention will improve long-term outcomes everywhere.
If your team wants a fast way to stress-test whether a methods section, evidence summary, or guideline sentence has drifted away from its actual PICO, Aqrab is built for that kind of critique. The simplest route is to start in Aqrab Try and force the population, comparator, and outcome assumptions into the open before they harden into a recommendation.
The Bottom Line
Indirectness does not mean the study was badly done. It means the evidence and the decision are no longer perfectly aligned. That mismatch matters because a clean estimate cannot rescue a question drift it was never designed to answer.
The sentence worth keeping is this: good evidence can still be the wrong evidence for the decision in front of you.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Guideline Recommendation Strength: When “Strongly Recommend” Starts Outrunning the Evidence
A practical guide to recommendation strength in clinical guidelines. Covers certainty of evidence, benefit-harm tradeoffs, patient values, implementation burden, and what reviewers should demand before trusting a forceful recommendation.
PROBAST: When a Prediction Model Paper Looks Ready Before It Earns Trust
A practical guide to PROBAST for clinical researchers. Covers participant selection, predictor leakage, outcome definition, overfitting, calibration, and what reviewers should demand before trusting a clinical prediction model.
Transitivity in Network Meta-Analysis: When Indirect Comparisons Pretend the Trials Were Exchangeable
A practical guide to transitivity in network meta-analysis for clinical researchers. Covers effect modifiers, shared comparators, indirect comparison failure modes, and what reviewers should demand before trusting rankings.