← Back to Blog
Clinical TrialsEvidence AppraisalMethods Critique

Clinical Importance: When a Significant Result Is Still Too Small to Matter

September 23, 2026·12 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A trial reports a mean difference of 3 points with a 95% confidence interval from 1 to 4 and p < 0.05. The result is statistically significant. If patients would need at least a 5-point between-group benefit to justify treatment burden, the same interval says something less celebratory: every effect compatible with the interval falls below the chosen threshold.

The null value and the clinical-importance threshold answer different questions. Reviewers should put both on the same scale before deciding what the result means.

The Method in One Sentence

Prespecify and justify an effect threshold for the treatment comparison, then read the full confidence interval against both zero and that threshold instead of allowing the p-value to make the decision.

Concrete takeaway

Draw two vertical lines on the effect plot: one at no effect and one at the smallest between-group benefit judged important. The confidence interval—not the significance label—shows which interpretations remain compatible with the data and model.

One Estimate, Two Reference Lines

Suppose higher scores are better and the prespecified smallest important between-group benefit is +5 points. The examples below are hypothetical; they teach interpretation, not a clinical cutoff.

PatternEstimate95% CIWhat the interval says
Clearly important+8+6 to +10The entire interval is beyond +5.
Significant, importance uncertain+7+2 to +12The interval excludes 0 but crosses +5.
Inconclusive+3−1 to +7The interval crosses both 0 and +5.
Precisely below the threshold+3+1 to +4The interval excludes 0 but stays below +5.

Assumptions: positive values favor treatment; +5 is a justified threshold for the between-arm contrast; the estimand, model, outcome scale, and follow-up are otherwise appropriate.

“Significant” Can Still Be Too Small

A significance test usually asks how incompatible the observed data are with a null hypothesis under specified assumptions. It does not establish that the effect is large enough to matter, that patients value the outcome, or that benefits outweigh harms and burden.

In the fourth row, the interval excludes zero, so a conventional two-sided test would reject no mean difference at the 5% level. Yet the upper confidence limit remains below +5. Under the stated assumptions, the result is precise enough to exclude the prespecified benefit threshold. The evidence supports a small positive effect, not an important one by that rule.

“Not Significant” Can Still Be Too Uncertain

In the third row, the interval includes zero and effects above +5. Calling the result “negative” loses both facts. The study did not demonstrate benefit, but it also did not rule out an important benefit. The useful diagnosis is imprecision, not proof of no effect.

This distinction affects the next decision. A precise estimate below the threshold can deprioritize a treatment for that outcome. A wide interval spanning harm, no effect, and important benefit signals that the evidence cannot settle the question.

Do Not Borrow a Threshold From the Wrong Level

A meaningful within-patient change is not automatically the smallest important difference between randomized groups. One describes how much change matters for an individual patient; the other describes the contrast in average outcomes or another prespecified treatment-effect measure between groups. FDA's patient-focused guidance explicitly separates these concepts.

For example, a treatment may shift the distribution modestly while increasing the proportion of patients who achieve a meaningful improvement. Conversely, a mean difference that equals an individual change threshold does not by itself show that the same number is the right decision boundary for a group comparison. Name the estimand and justify the threshold at that level.

A Threshold Is a Judgment, Not a Biological Constant

The smallest important effect can depend on adverse effects, treatment burden, cost, available alternatives, disease severity, duration of benefit, and whose preferences count. A threshold copied from another population or outcome version may not transport.

Use patient input, anchor-based evidence, clinical judgment, prior trials, and decision context as appropriate. Prefer a justified range when one exact value would imply false certainty. Most importantly, set the rule before seeing the estimate; choosing the threshold after the result invites a convenient conclusion.

Do Not Turn This Into an Equivalence Claim

If the upper confidence limit is below a positive benefit threshold, the analysis may rule out benefit of that magnitude for the stated estimand. That does not automatically prove that treatments are equivalent, interchangeable, or similarly safe. Equivalence requires a design and analysis with justified two-sided margins, appropriate operating characteristics, and attention to assay sensitivity and adherence.

Keep the conclusion narrow: “The interval excludes benefits of 5 points or more on this outcome over 12 weeks.” Then appraise harms, secondary outcomes, missing data, multiplicity, and generalizability separately.

The 90-Second Clinical-Importance Audit

  • Was the threshold defined before the result was seen?
  • Does it represent a between-group effect or a within-patient change—and are those being confused?
  • Who judged the effect important: patients, clinicians, regulators, or investigators?
  • Is the threshold justified for this outcome, population, treatment burden, and time horizon?
  • Does the confidence interval exclude the null, cross the threshold, or sit entirely below it?
  • Are benefits, harms, inconvenience, cost, and uncertainty considered together?

What Reviewers Should Demand

Ask for the prespecified treatment-effect measure, point estimate, confidence interval, null value, and clinical-importance threshold on one display. The protocol or statistical analysis plan should explain where the threshold came from and which population, outcome, and time horizon it represents.

Then use language that matches the interval. “Statistically significant” is not a synonym for “clinically important.” “Not significant” is not a synonym for “no effect.” A conclusion should say whether important benefit is supported, excluded, or still compatible with the evidence.

Sources and Evidence Maturity

  • ICH E9: Statistical Principles for Clinical Trials — final international guideline; states that trial objectives should specify the treatment effect of interest, clinically relevant differences should be considered in planning, and estimates should be accompanied by confidence intervals where possible.
  • FDA Patient-Focused Drug Development Guidance 3 — final guidance issued October 2025; addresses fit-for-purpose clinical outcome assessments and evidence that an assessment is understandable and relevant in its context of use.
  • FDA Patient-Focused Drug Development Guidance Series — official program page; identifies Guidance 4 as draft guidance on COA-based endpoints, meaningful change, and interpretation. Draft recommendations are not final agency policy.

Evidence note: the +5 threshold and all four numerical examples are hypothetical. A real threshold requires outcome-, population-, estimand-, and decision-specific justification.

Where Aqrab Fits

Aqrab can help reviewers extract the estimand, outcome scale, effect estimate, confidence interval, prespecified threshold, follow-up, analysis population, and source of threshold justification across a protocol and paper. Use that structured comparison to expose a significance-only conclusion—not to invent a universal cutoff or replace clinical judgment. Try Aqrab on a trial report, or explore plans for repeatable evidence-review workflows.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive
Next guide

This is the newest guide so far.