Clinical Importance: When a Significant Result Is Still Too Small to Matter
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
A trial reports a mean difference of 3 points with a 95% confidence interval from 1 to 4 and p < 0.05. The result is statistically significant. If patients would need at least a 5-point between-group benefit to justify treatment burden, the same interval says something less celebratory: every effect compatible with the interval falls below the chosen threshold.
The null value and the clinical-importance threshold answer different questions. Reviewers should put both on the same scale before deciding what the result means.
The Method in One Sentence
Prespecify and justify an effect threshold for the treatment comparison, then read the full confidence interval against both zero and that threshold instead of allowing the p-value to make the decision.
Concrete takeaway
Draw two vertical lines on the effect plot: one at no effect and one at the smallest between-group benefit judged important. The confidence interval—not the significance label—shows which interpretations remain compatible with the data and model.
One Estimate, Two Reference Lines
Suppose higher scores are better and the prespecified smallest important between-group benefit is +5 points. The examples below are hypothetical; they teach interpretation, not a clinical cutoff.
| Pattern | Estimate | 95% CI | What the interval says |
|---|---|---|---|
| Clearly important | +8 | +6 to +10 | The entire interval is beyond +5. |
| Significant, importance uncertain | +7 | +2 to +12 | The interval excludes 0 but crosses +5. |
| Inconclusive | +3 | −1 to +7 | The interval crosses both 0 and +5. |
| Precisely below the threshold | +3 | +1 to +4 | The interval excludes 0 but stays below +5. |
Assumptions: positive values favor treatment; +5 is a justified threshold for the between-arm contrast; the estimand, model, outcome scale, and follow-up are otherwise appropriate.
“Significant” Can Still Be Too Small
A significance test usually asks how incompatible the observed data are with a null hypothesis under specified assumptions. It does not establish that the effect is large enough to matter, that patients value the outcome, or that benefits outweigh harms and burden.
In the fourth row, the interval excludes zero, so a conventional two-sided test would reject no mean difference at the 5% level. Yet the upper confidence limit remains below +5. Under the stated assumptions, the result is precise enough to exclude the prespecified benefit threshold. The evidence supports a small positive effect, not an important one by that rule.
“Not Significant” Can Still Be Too Uncertain
In the third row, the interval includes zero and effects above +5. Calling the result “negative” loses both facts. The study did not demonstrate benefit, but it also did not rule out an important benefit. The useful diagnosis is imprecision, not proof of no effect.
This distinction affects the next decision. A precise estimate below the threshold can deprioritize a treatment for that outcome. A wide interval spanning harm, no effect, and important benefit signals that the evidence cannot settle the question.
Do Not Borrow a Threshold From the Wrong Level
A meaningful within-patient change is not automatically the smallest important difference between randomized groups. One describes how much change matters for an individual patient; the other describes the contrast in average outcomes or another prespecified treatment-effect measure between groups. FDA's patient-focused guidance explicitly separates these concepts.
For example, a treatment may shift the distribution modestly while increasing the proportion of patients who achieve a meaningful improvement. Conversely, a mean difference that equals an individual change threshold does not by itself show that the same number is the right decision boundary for a group comparison. Name the estimand and justify the threshold at that level.
A Threshold Is a Judgment, Not a Biological Constant
The smallest important effect can depend on adverse effects, treatment burden, cost, available alternatives, disease severity, duration of benefit, and whose preferences count. A threshold copied from another population or outcome version may not transport.
Use patient input, anchor-based evidence, clinical judgment, prior trials, and decision context as appropriate. Prefer a justified range when one exact value would imply false certainty. Most importantly, set the rule before seeing the estimate; choosing the threshold after the result invites a convenient conclusion.
Do Not Turn This Into an Equivalence Claim
If the upper confidence limit is below a positive benefit threshold, the analysis may rule out benefit of that magnitude for the stated estimand. That does not automatically prove that treatments are equivalent, interchangeable, or similarly safe. Equivalence requires a design and analysis with justified two-sided margins, appropriate operating characteristics, and attention to assay sensitivity and adherence.
Keep the conclusion narrow: “The interval excludes benefits of 5 points or more on this outcome over 12 weeks.” Then appraise harms, secondary outcomes, missing data, multiplicity, and generalizability separately.
The 90-Second Clinical-Importance Audit
- Was the threshold defined before the result was seen?
- Does it represent a between-group effect or a within-patient change—and are those being confused?
- Who judged the effect important: patients, clinicians, regulators, or investigators?
- Is the threshold justified for this outcome, population, treatment burden, and time horizon?
- Does the confidence interval exclude the null, cross the threshold, or sit entirely below it?
- Are benefits, harms, inconvenience, cost, and uncertainty considered together?
What Reviewers Should Demand
Ask for the prespecified treatment-effect measure, point estimate, confidence interval, null value, and clinical-importance threshold on one display. The protocol or statistical analysis plan should explain where the threshold came from and which population, outcome, and time horizon it represents.
Then use language that matches the interval. “Statistically significant” is not a synonym for “clinically important.” “Not significant” is not a synonym for “no effect.” A conclusion should say whether important benefit is supported, excluded, or still compatible with the evidence.
Sources and Evidence Maturity
- ICH E9: Statistical Principles for Clinical Trials — final international guideline; states that trial objectives should specify the treatment effect of interest, clinically relevant differences should be considered in planning, and estimates should be accompanied by confidence intervals where possible.
- FDA Patient-Focused Drug Development Guidance 3 — final guidance issued October 2025; addresses fit-for-purpose clinical outcome assessments and evidence that an assessment is understandable and relevant in its context of use.
- FDA Patient-Focused Drug Development Guidance Series — official program page; identifies Guidance 4 as draft guidance on COA-based endpoints, meaningful change, and interpretation. Draft recommendations are not final agency policy.
Evidence note: the +5 threshold and all four numerical examples are hypothetical. A real threshold requires outcome-, population-, estimand-, and decision-specific justification.
Where Aqrab Fits
Aqrab can help reviewers extract the estimand, outcome scale, effect estimate, confidence interval, prespecified threshold, follow-up, analysis population, and source of threshold justification across a protocol and paper. Use that structured comparison to expose a significance-only conclusion—not to invent a universal cutoff or replace clinical judgment. Try Aqrab on a trial report, or explore plans for repeatable evidence-review workflows.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
CONSORT 2025 for Reviewers: A Practical Checklist for Interpretable Trials
A practical CONSORT 2025 reviewer checklist for randomized trials. Find the reporting gaps that change what you can infer about allocation, treatment, outcomes, missing data, and analysis.
Fragility Index: When One or Two Events Carry More Confidence Than They Should
A practical guide to the fragility index for clinical researchers. Covers event-flip sensitivity, loss to follow-up, effect-size context, and what reviewers should demand before trusting a barely significant trial.
Absolute Effects: Why the Same Relative Risk Can Mean 5 or 100 Fewer Events
A practical guide to translating relative effects into absolute effects. Anchor the estimate to a credible comparator risk, keep the time horizon visible, and show uncertainty before quoting an NNT.
This is the newest guide so far.