← Back to Blog
Clinical TrialsMultiplicityMethods Critique

Hierarchical Testing in Clinical Trials: When a Significant Secondary Endpoint Is Still Descriptive

September 10, 2026·12 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A trial misses its primary endpoint at p = 0.08. A secondary endpoint reports p = 0.01. The abstract celebrates the secondary result. Is that result confirmatory? The number alone cannot answer. You need the prespecified multiple-testing procedure—and you need to know whether the procedure ever reached that endpoint.

Hierarchical testing is often described as a ranked list. It is more precise to treat it as an algorithm for spending a trial's Type I error budget. If an earlier gate does not open, a later nominal p-value can remain scientifically interesting without supporting the confirmatory claim the headline implies.

The Design in One Sentence

In a fixed-sequence procedure, hypotheses are ordered in advance and tested one at a time at a prespecified alpha level; testing proceeds only while each required earlier hypothesis is rejected.

Concrete takeaway

Reconstruct the testing path before interpreting any secondary p-value. “Statistically significant” and “formally tested under multiplicity control” are not synonyms.

A Hierarchy Is a Path, Not a Decoration

Gate 1

Primary endpoint

Test at the prespecified level. If it passes, continue.

Gate 2

Key secondary A

Formally test only if Gate 1 released the path.

Gate 3

Key secondary B

A small nominal p-value does not reopen a closed path.

The order should reflect the scientific claim structure, not the final pattern of results. FDA guidance emphasizes prospective specification of endpoint families, analyses, and the testing procedure. ICH E9 similarly separates pre-defined confirmatory hypotheses from exploratory analyses. An order invented after unblinding describes the data; it does not restore the error control of an order chosen before seeing them.

Why the Gate Exists

Every additional opportunity to declare success can increase the chance of at least one false-positive conclusion. The relevant problem is not that a secondary endpoint is inherently less worthy. It is that the family of planned claims creates multiple chances to be wrong. A valid procedure limits that overall chance while defining which conclusions the trial may support.

For the broader problem—including repeated time points, subgroup searches, interim looks, and dose choices—start with Aqrab's guide to multiple testing in clinical trials. This guide focuses on the narrower reviewer task that follows: replaying a prespecified endpoint hierarchy gate by gate.

Fixed-sequence testing can be efficient when endpoints have a defensible order: no alpha is divided among later tests while the path remains open. The tradeoff is fragility. One failed gate can stop formal testing of all endpoints below it, even when a later effect is large or clinically important. That later estimate should still be reported with its uncertainty; its status becomes descriptive or exploratory under that particular procedure.

Three Questions That Look Similar but Are Not

QuestionWhat answers itCommon error
Was there an observed difference?Effect estimate, interval, and data qualityEquating a small p-value with importance
Was the endpoint confirmatory?Prespecified family and testing pathIgnoring a closed gate
Is the effect clinically meaningful?Endpoint validity, magnitude, uncertainty, and contextTreating error control as clinical interpretation

Multiplicity control answers a statistical claim question. It does not make a trivial difference important, validate a poor outcome measure, repair missing data, or guarantee that the estimate will transport to routine care.

Worked Example: The Endpoint That Never Got Tested

Suppose a protocol specifies a fixed sequence: functional status first, symptom burden second, hospitalization third. Each hypothesis may be tested at two-sided alpha 0.05 only if the previous one succeeds. Results are p = 0.08, 0.004, and 0.03, respectively.

The procedure stops at functional status. The symptom and hospitalization results may be important signals, but their nominal p-values are not confirmatory results from the fixed sequence. Reporting “symptoms improved significantly” without that qualification erases the testing rule. A defensible summary reports the missed primary endpoint, the effect estimates and intervals for all outcomes, and the exploratory status of later findings.

Change the procedure and the answer can change. A graphical or parallel gatekeeping plan may distribute and recycle alpha across branches. That is why “the primary failed, so every secondary is invalid” is also too crude. The exact prespecified algorithm—not a universal slogan—sets the claim boundary.

The 90-Second Hierarchical Testing Audit

  • Which hypotheses belonged to the confirmatory family?
  • Was the testing order fixed before treatment assignments were unmasked?
  • How much alpha was available at each step, and what event released it?
  • Did a failed hypothesis close the path to later endpoints?
  • Are reported p-values adjusted, nominal, or purely descriptive?
  • Does the conclusion match the last hypothesis reached by the procedure?

What Reviewers Should Demand

Ask for the protocol and statistical analysis plan version finalized before unblinding. Map every confirmatory hypothesis to its endpoint definition, estimand, time point, population, model, and alpha. Then replay the procedure using the reported results. A table of raw p-values is not a multiplicity strategy, and the label “key secondary” does not show that a hypothesis inherited alpha.

If the publication departs from the plan, the change may have a sound operational explanation. It still needs a date, rationale, and transparent effect on interpretation. Preserve useful exploratory findings, but label them honestly and treat them as hypotheses for confirmation rather than retroactive proof.

Sources and Evidence Maturity

Evidence note: this guide interprets mature statistical and regulatory guidance. It does not assess a specific treatment or make a clinical effectiveness claim. The numerical example is hypothetical.

Where Aqrab Fits

Aqrab can help make an endpoint hierarchy inspectable. Bring the protocol, statistical analysis plan, and results together, then ask Aqrab to extract the confirmatory family, testing order, alpha transitions, stopping rules, and the highest claim each reported result supports. Use the output as a structured review aid—not a substitute for the trial statistician or the source documents. Explore Aqrab plans for repeatable evidence-review workflows.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive