Randomized Withdrawal Trials: Why a Relapse-Prevention Win Is Not a New-Patient Effect
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
A trial gives every participant the active treatment. People who improve and tolerate it continue; everyone else exits before randomization. The apparent responders are then assigned either to keep treatment or switch to placebo. Relapse is less common among those who continue. The treatment has demonstrated a maintenance effect in selected responders—not an average benefit for every patient who might start it.
That distinction is the center of a randomized withdrawal trial. The design can answer an important question efficiently and may limit prolonged placebo exposure. Its strength comes from enrichment. Its main interpretive hazard comes from forgetting what that enrichment removed.

The Design in One Sentence
Among participants who showed an apparent response—often also adequate tolerability—during an initial treatment period, randomization compares continuing treatment with withdrawing it, commonly by switching to placebo and monitoring relapse or loss of control.
Concrete takeaway
Read the eligibility criteria at randomization, not only at initial enrollment. The causal contrast belongs to the people who passed both gates.
Two Cohorts Exist, but Only One Is Randomized
| Stage | Who is represented | What can be learned |
|---|---|---|
| Initial run-in | All screened participants who start active treatment | Observed response, tolerability, adherence, and attrition under a usually nonrandomized phase |
| Randomized withdrawal | The subset meeting prespecified continuation criteria | Effect of continuing versus withdrawing under the protocol's follow-up and rescue rules |
| Target practice population | New patients considering treatment | Not directly estimated unless uptake, early failure, harms, and representativeness are also addressed |
Randomization protects the post-run-in comparison from baseline confounding within the selected group. It does not randomize the selection process that created that group. The proportion who entered, responded, tolerated treatment, adhered, and reached randomization is therefore part of the result—not a disposable screening detail.
What the Trial Can Claim
A well-conducted design can support a claim that continued treatment preserves disease control better than withdrawal over the studied interval among eligible apparent responders. The FDA describes randomized withdrawal as an empirical enrichment strategy and notes its value when prolonged placebo treatment would be ethically or practically difficult.
The design may also be sensitive to ongoing pharmacologic activity: if withdrawing treatment leads to prespecified recurrence while continued treatment maintains control, the contrast is difficult to explain by stable baseline prognosis alone. Blinding, a credible withdrawal process, and objective relapse definitions still matter because expectations, rescue decisions, and withdrawal symptoms can influence the observed endpoint.
What the Trial Cannot Claim by Itself
Average benefit when treatment is first started
Nonresponders and early intolerance are filtered before randomization. A strong randomized effect among responders does not supply an intention-to-treat effect for all starters.
Benefit in patients unlike the enriched cohort
Run-in rules may favor adherent, tolerant, lower-risk, or otherwise selected participants. Generalization requires a comparison between randomized responders and the clinical population of interest.
Complete long-term safety
Removing people with early adverse effects and randomizing a smaller responder subset can reduce the information available for population-level harms, rare events, and tolerability.
The Withdrawal Arm Is an Intervention Too
“Placebo” may conceal clinically important details. Was active therapy stopped abruptly or tapered? Could abrupt stopping cause rebound or withdrawal symptoms that resemble relapse? How quickly could rescue treatment begin? Were outcome assessors blinded to symptoms that might reveal assignment? These design choices determine whether the comparison reflects maintenance of benefit, consequences of stopping, or both.
ICH E9(R1) requires the treatment conditions, target population, outcome, summary measure, and handling of intercurrent events to be explicit. In a withdrawal trial, rescue medication, restart of active therapy, discontinuation, and loss to follow-up can each alter what the endpoint means. Labeling every post-randomization complication as “missing data” hides the clinical question.
The 90-Second Randomized Withdrawal Audit
- Who entered the run-in, and how many never reached randomization?
- How were response and tolerability defined, measured, and timed?
- What exactly differed after randomization: continuation, abrupt withdrawal, taper, or rescue strategy?
- Was relapse a clinically meaningful outcome with blinded, prespecified ascertainment?
- Were withdrawal symptoms distinguishable from recurrence of the disease?
- Did analyses and conclusions stay within the randomized responder population?
Worked Example: A Strong Maintenance Result With a Narrow Population
Imagine 1,000 patients begin open-label treatment. Two hundred stop for adverse effects, 350 do not meet the response threshold, and 50 leave for other reasons. The remaining 400 apparent responders are randomized. By week 16, relapse occurs in 12% continuing treatment and 34% switched to placebo.
The randomized risk difference is 22 percentage points among those 400 responders under the trial's withdrawal and rescue rules. That is a clinically useful maintenance contrast. It does not mean treatment prevents relapse in 22% of all 1,000 starters. The 600 people who did not reach randomization remain central to any claim about beginning therapy in routine care.
Sources and Evidence Maturity
- FDA, Enrichment Strategies for Clinical Trials to Support Approval of Human Drugs and Biological Products (March 2019) — final regulatory guidance; primary source for the definition, uses, strengths, and limitations of randomized withdrawal enrichment.
- ICH E9(R1), Addendum on Estimands and Sensitivity Analysis in Clinical Trials (2019) — Step 4 international harmonized guideline; primary source for aligning the target population, treatment conditions, outcome, intercurrent-event strategy, and trial design.
Evidence note: these are mature design and regulatory sources, not comparative patient-outcome studies. The worked numbers are hypothetical and demonstrate interpretation, not an empirical treatment effect.
Where Aqrab Fits
Aqrab helps reviewers separate the population that started treatment from the population that was randomized. Paste a trial's flow diagram, eligibility criteria, and endpoint definition into Aqrab Try and ask it to extract both gates, the withdrawal strategy, rescue rules, and the defensible claim ceiling. For repeatable evidence review, use the developer documentation to add the audit to a structured workflow.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Hierarchical Testing in Clinical Trials: When a Significant Secondary Endpoint Is Still Descriptive
A practical hierarchical testing guide for clinical researchers. Reconstruct the prespecified testing path before treating a small p-value on a secondary endpoint as confirmatory evidence.
Allocation Concealment: The Randomized-Trial Safeguard That Works Before Assignment
A practical allocation concealment guide for clinical researchers and peer reviewers. Separate sequence generation, concealment, implementation, and blinding before trusting the word randomized.
Cluster-Trial Recruitment Bias: When the Clinic Knows the Assignment Before the Patient Enters
A practical guide to post-randomization recruitment bias in cluster trials. Audit timing, allocation awareness, eligibility, consent, patient denominators, and the limits of statistical adjustment.