← Back to Blog
Clinical TrialsEquivalenceMethods Critique

Equivalence Trials: Why ‘Not Significant’ Does Not Mean ‘The Same’

September 22, 2026·12 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Two treatments produce a non-significant difference. The abstract calls them equivalent. That conclusion may be exactly backward: a wide, uninformative confidence interval often crosses zero because the study could not distinguish a meaningful benefit from a meaningful harm.

Equivalence requires a different question, a prespecified zone of clinically unimportant differences, and evidence precise enough to place the whole compatible range inside that zone.

Two confidence-interval plots: one wide interval crosses zero and both equivalence margins, while a second narrow interval lies completely inside the margins
Crossing the no-difference line answers neither side of the equivalence question. Precision relative to both margins is what matters.

The Method in One Sentence

Choose lower and upper differences that would still be clinically unimportant, then conclude equivalence only when the confidence interval for the treatment contrast lies entirely between those prespecified margins.

Concrete takeaway

A superiority test asks whether the effect differs from zero. An equivalence test asks whether effects at least as large as either clinically important boundary can be excluded. Failing the first test does not pass the second.

Draw Three Lines Before Looking at the Data

Put the no-difference value in the center. Then draw a lower margin, −Δ, and an upper margin, +Δ. These boundaries should represent the largest disadvantages in either direction that clinicians, patients, and regulators would still consider unimportant for the intended use.

The margin is not a number chosen to make the result fit. Its scale, direction, and clinical rationale belong in the protocol. A risk difference of ±5 percentage points, a mean difference of ±2 scale units, and a ratio range around 1 encode different clinical judgments. A convenient historical convention is not automatically appropriate for a new endpoint or setting.

Two One-Sided Questions

The two one-sided tests procedure asks whether the new treatment is above the lower boundary and below the upper boundary. Both null hypotheses must be rejected. With two one-sided tests at the 5% level, the corresponding 90% confidence interval must fit fully inside the equivalence interval.

Confidence intervalDefensible reading
Entirely inside both marginsSupports equivalence on the prespecified scale
Crosses zero and one or both marginsInconclusive, not equivalent
Entirely outside zero but inside both marginsStatistically different yet clinically equivalent
Inside one margin but beyond the otherMay support one-sided noninferiority, not equivalence

Worked Example: The Same Point Estimate, Opposite Conclusions

Suppose lower symptom scores are better and the equivalence margins are −3 and +3 points. Trial A estimates a difference of 0.4 with a 90% confidence interval from −5.2 to +6.0. Its ordinary superiority p-value is non-significant, but clinically important benefit and harm remain compatible with the data. The result is inconclusive.

Trial B also estimates 0.4, with a 90% confidence interval from −1.1 to +1.9. That interval sits inside both margins, so it supports equivalence. The point estimates are identical; the inferential claims differ because precision differs.

Why Bias Toward Similarity Is Dangerous

In superiority trials, dilution often makes an effect harder to detect. In equivalence trials, the same dilution can make treatments look reassuringly alike. Poor adherence, treatment crossover, insensitive outcome measurement, protocol deviations, and missing outcomes can all compress observed differences.

That is why reporting both intention-to-treat and per-protocol analyses is especially informative here. Agreement strengthens interpretation; disagreement is a warning that the equivalence claim depends on analysis-set choices. Neither analysis repairs a badly conducted study, and a per-protocol subset is not protected by randomization in the same way as the original assignment.

Equivalence Is Not Noninferiority

Noninferiority excludes an unacceptable disadvantage in one direction. Equivalence excludes clinically important differences in both directions. A product may be no worse than a comparator by more than the lower margin while still being meaningfully better—or while the upper boundary remains poorly estimated. Calling that result equivalent erases the second half of the question.

Bioequivalence is a specialized regulatory application, commonly evaluated on log-transformed pharmacokinetic measures and expressed through ratios. Its familiar acceptance ranges should not be copied into clinical equivalence trials without endpoint-specific justification.

The 90-Second Equivalence Audit

  • Was equivalence the prespecified primary objective?
  • Are both lower and upper equivalence margins stated and clinically justified?
  • Is the analysis scale—difference, ratio, or transformed ratio—explicit?
  • Does the entire confidence interval lie inside both margins?
  • Are intention-to-treat and per-protocol results both reported and interpreted?
  • Could poor adherence, missing data, or measurement noise have pulled the groups artificially together?

What Reviewers Should Demand

Require the equivalence objective and both margins to appear in the protocol, registration, sample-size calculation, analysis plan, and report. Ask who judged the margins clinically unimportant, what evidence informed them, and whether the confidence level matches the testing procedure.

Then inspect the confidence interval against both boundaries, not merely the p-value against zero. Review adherence, crossover, missingness, assay sensitivity, and analysis-set agreement. The correct conclusion may be superiority, equivalence, noninferiority only, or simple uncertainty; the plot should make those possibilities visible.

Sources and Evidence Maturity

Evidence note: this guide synthesizes mature regulatory and reporting standards. The symptom-score example is hypothetical and does not establish a universal margin.

Where Aqrab Fits

Aqrab can help reviewers locate the objective, margin rationale, analysis populations, confidence intervals, deviations, and missing-data handling across a protocol and report. Use that structure to expose claim drift—not to replace clinical judgment about what difference is unimportant. Try Aqrab on a trial report, or explore plans for repeatable evidence-review workflows.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive
Next guide

This is the newest guide so far.