← Back to Blog
Clinical TrialsEvidence AppraisalMethods Critique

Fragility Index: When One or Two Events Carry More Confidence Than They Should

June 12, 2026·15 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Some randomized trials earn their abstract confidence by a margin that is almost comically thin. Change one or two events in the better-looking arm and the result loses nominal significance. That does not prove the treatment fails. It does tell you the rhetorical posture of the paper may be sturdier than the event table underneath it.

The fragility index is a simple stress test for binary outcomes. It asks how many event status changes would be needed to move a statistically significant comparison above the conventional 0.05 threshold. Used well, it helps clinical readers calibrate certainty. Used badly, it becomes another single-number fetish.

What the Fragility Index Is Actually Measuring

Start with a two-arm trial and a binary endpoint: death, stroke, relapse, hospitalization, healing, or another yes-or-no outcome. If the intervention arm looks better and the p-value is below 0.05, ask how many non-events would have to become events in that better-looking arm before the result stops crossing the significance line.

A fragility index of 1 means one event flip is enough. A fragility index of 8 means the result survives more perturbation. Neither number tells you whether the study is unbiased, clinically important, or generalizable. It only tells you how tightly the inferential badge is tied to a small count of outcome changes.

Why This Matters for Trial Reading

Tiny event counts carry huge rhetorical weight

When only a few additional events would change the conclusion, the difference between “positive” and “not significant” may be thinner than the abstract suggests.

Loss to follow-up can dominate fragility

If more patients are missing outcome data than the number of flips needed to erase significance, the trial deserves a more cautious read.

Nominal significance is not robustness

A p-value can be below 0.05 and still rest on a result that is operationally brittle, especially in smaller trials with few total events.

Interactive fragility-index explorer

Change a few outcome events and watch a statistically significant trial lose its badge

This calculator uses a two-sided Fisher exact test for a simple binary outcome table. The fragility index is the number of additional events needed in the better-looking group to move the result above 0.05.

Fragility index1event flips needed to lose nominal significance

This lets you compare outcome fragility against missing-outcome uncertainty. Many flashy trials have more patients lost to follow-up than event flips required to lose significance.

Two-sided p-value

0.035

Nominal significance is only part of the story.

Treated risk

8.3%

Observed event rate in the intervention arm.

Control risk

18.3%

Observed event rate in the comparator arm.

Lost to follow-up / fragility

9.0

Ratios above 1 should make the interpretation more careful than the abstract usually is.

QuantityValueInterpretation
Absolute risk difference-10.0%Effect size still matters. A low fragility index is not a synonym for a trivial effect or a false one.
Outcome flips needed1How many extra events in the better-looking group would move the Fisher exact p-value to 0.05 or higher.
Lost to follow-up9Missing outcomes can outweigh nominal event robustness, especially in smaller pragmatic trials.
Decision cueHighly fragileUse this as an interpretation prompt, not a replacement for effect size, bias assessment, or trial design judgment.

Where the Metric Helps

The fragility index is useful when readers are being nudged toward false dichotomies. A small trial may report a dramatic relative risk reduction and a triumphant p-value, but the underlying event table may be one flipped hospitalization away from losing its positive label. That is worth knowing before the result gets translated into guideline confidence or product-copy certainty.

It is especially helpful beside the number lost to follow-up. If a trial's fragility index is 2 but 11 patients have missing primary-outcome status, then the outcome uncertainty already exceeds the event changes required to erase nominal significance. That does not invalidate the effect. It does weaken any swagger built on the p-value alone.

Where the Metric Misleads

Do not turn it into a replacement for effect size

A low fragility index can coexist with an important effect in a small high-risk population. A higher index can coexist with a clinically trivial effect. Magnitude still matters.

Do not ignore bias, ascertainment, or adjudication

Event fragility is only one vulnerability. Outcome misclassification, unblinded adjudication, selective endpoint choice, and differential surveillance can matter more.

Do not overextend beyond simple binary comparisons

The classic fragility index is most natural for parallel-group trials with binary outcomes. Time-to-event analyses and adjusted models need more careful analogues.

Do not act as if 0.05 is a law of nature

The metric inherits the same threshold dependence as the p-value it is interrogating. Its value is in exposing threshold fragility, not in sanctifying the threshold.

A Concrete Clinical Example

Case

A pragmatic heart-failure readmission trial with a low event count

Imagine a pragmatic trial of a post-discharge support program. The intervention arm has fewer 30-day readmissions, the p-value slips under 0.05, and the headline celebrates a positive result. But the fragility index is 2 and seven patients are lost to follow-up or have uncertain readmission capture because they used outside hospitals.

The right interpretation is not “the study is worthless.” It is “the result may be promising, but it is event-fragile and missing-outcome uncertainty is larger than the margin of nominal significance.” That is a very different conclusion from “this intervention is clearly effective.”

Reviewer Decision Table

Observed patternWhat it meansBetter interpretation move
Fragility index 1 to 3, few total eventsNominal significance depends on a tiny number of outcomesTemper certainty, emphasize confidence intervals, and inspect event adjudication closely.
Loss to follow-up exceeds fragility indexMissing outcomes may outweigh observed inferential marginAsk how outcome status for missing participants could change the conclusion.
Larger fragility index but tiny absolute effectResult may be statistically sturdier than it is clinically importantKeep clinical magnitude and patient relevance in the foreground.
Adjusted or time-to-event analysis onlyClassic fragility index may not map cleanly to the estimandUse the idea as a caution prompt, not as a ritualized scorecard.

What Aqrab Should Help a Team Do Here

This is the kind of evidence-appraisal judgment that gets flattened when teams reduce trial reading to a p-value and a quote from the abstract. The better workflow is to pair the event table with missingness, adjudication, endpoint choice, and practical consequence before declaring the result solid.

If your team wants a faster way to stress-test trial claims, reviewer red flags, and interpretation drift before those claims harden into slides or manuscripts, Aqrab's critique workflow is built for that kind of methodological reading.

Bottom Line

The fragility index is not a truth detector. It is a reminder that “statistically significant” can rest on a surprisingly small number of events. Use it to discipline interpretation, especially when the study is small, follow-up is incomplete, and the abstract sounds more decisive than the data table deserves.