← Back to Blog
Clinical TrialsOutcome MeasurementMethods Critique

Baseline Adjustment in Randomized Trials: Why Change From Baseline Keeps Losing to ANCOVA

July 10, 2026·15 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Randomization does not abolish baseline values. It simply means the baseline values should be handled intelligently instead of ritualistically.

Yet trial reports still fall in love with change from baseline, percent change, and responder cutoffs as if subtraction automatically makes an analysis fairer. Often it just makes it noisier. Sometimes it makes it sillier.

The Core Design Rule

For a continuous outcome measured once after randomization, the default primary analysis is usually an adjusted follow-up model such as ANCOVA: compare groups on the follow-up value while adjusting for the baseline measure and any prespecified stratification factors.

Decision rule:

If a manuscript uses change from baseline or percent change as the primary endpoint, ask what that analysis does better than ANCOVA. If the answer is mainly "it felt intuitive," the method is already in trouble.

The reason is not fashion. Adjusted follow-up models usually deliver better precision, cope better with chance baseline imbalance, and keep the estimand focused on the post-randomization outcome that actually matters.

Why Change From Baseline Feels So Reasonable

It sounds patient-centered

Researchers like saying patients improved by 8 points rather than ended at a mean of 14. That is a communication preference, not an argument for the primary model.

It seems to handle baseline automatically

Subtracting baseline feels like control, but it bundles baseline measurement error directly into the outcome and often throws away efficiency in the process.

It flatters threshold headlines

Once the trial is written as change, it becomes easier to declare that 50% of patients achieved a clinically meaningful improvement. The marketing department applauds. The information content does not.

It survives by habit

Many teams use the same endpoint convention their field has always used. Familiarity is not a statistical justification. It is just older company.

What ANCOVA Is Actually Doing

ANCOVA does not punish patients for starting sicker. It uses the baseline value as prognostic information so the comparison at follow-up is less noisy. If baseline severity strongly predicts the outcome, adjusting for it usually narrows uncertainty. That is a feature, not suspicious efficiency.

It also responds better to the annoying but common fact that randomization in finite samples can still leave groups a bit imbalanced at baseline. Change scores do not magically erase that problem. They can actually magnify it when baseline values are noisy.

A Concrete Clinical Example

Imagine a moderate-size rheumatology trial comparing an anti-inflammatory therapy with placebo using a continuous disease-activity score at 12 weeks. Baseline severity predicts the week-12 score strongly, but the treated arm happens to start a little worse despite randomization.

Naive headline

The treated arm improved 9 points versus 6 points in controls, so the change score looks easy to sell and clinically intuitive.

What may be hiding

A worse starting point leaves more room to improve, baseline measurement error seeps directly into the change outcome, and percent change becomes unstable for low denominators.

What reviewers should ask

Show the ANCOVA estimate on the follow-up score, prespecify the threshold analyses as secondary, and explain whether missing follow-up could distort the polished change narrative.

Interactive baseline-analysis triage

Which analysis keeps the treatment effect and drops the avoidable statistical theater?

This is a teaching device, not a validated scoring system. Use it to see why adjusted follow-up models usually beat change scores, why percent change is a trouble magnet, and when threshold outcomes belong in the supporting cast.

Recommended defaultANCOVA / adjusted follow-up valueScore: 92/100

Higher correlation increases the efficiency gain from adjusting for baseline rather than subtracting it and hoping the noise cancels politely.

Small trials and noisy symptom scales make this more than a theoretical nuisance. ANCOVA handles it better than raw change scores.

This is where percent change starts behaving like a denominator prank rather than an outcome analysis.

A threshold can help communicate results. It should not usually replace the continuous primary analysis that actually preserved the information.

Precision tax of change scores

29 points

A rough reminder of how much efficiency you may be donating when an adjusted follow-up model was available.

Percent-change instability

12/100

When baseline denominators wobble, percent improvement can turn tiny numeric differences into fake drama.

Clinical-story distortion risk

21/100

The more imbalance and threshold pressure you have, the easier it becomes to oversell a clinically neat but statistically lossy headline.

Analysis optionTeaching scoreWhat it tends to buy you
ANCOVA / adjusted follow-up value92/100Best default for a single continuous follow-up outcome when you want efficiency without inventing extra noise.
Change from baseline63/100Can look intuitive, but usually pays a precision tax and handles imbalance less gracefully than adjusted follow-up models.
Percent change29/100Fragile when baseline values approach zero and too eager to turn denominators into drama.
Responder threshold22/100Useful mainly as a secondary clinical translation, not as the place where your information goes to die.

Comparison Matrix: What Each Analysis Usually Costs You

Primary analysis choiceWhen it can be reasonableMain failure mode
ANCOVA / adjusted follow-up valueDefault for a single continuous follow-up endpoint when baseline values are prognostic and measured before treatment.Usually not conceptual. The main risk is reporting it badly or skipping prespecified covariates.
Change from baselineSometimes acceptable for supportive presentation when paired with the adjusted primary result.Pays a precision penalty and can mishandle chance baseline imbalance while amplifying measurement error.
Percent changeRarely ideal; at best a carefully justified secondary descriptive view with stable denominators.Near-zero baseline values and denominator instability turn modest noise into theatrical swings.
Responder thresholdGood for clinical translation if the cutoff is prespecified, well justified, and secondary to a continuous primary analysis.Throws away information, invites threshold shopping, and can inherit baseline distortions.

Decision Rules for Trial Teams and Reviewers

  • Use an adjusted follow-up model as the primary analysis for a single continuous follow-up endpoint unless you have a strong protocol-level reason not to.
  • Treat change from baseline as a secondary presentation choice, not automatic methodological virtue.
  • Be deeply suspicious of percent change when baseline values can be small, noisy, or clinically heterogeneous.
  • Keep responder outcomes secondary unless the dichotomy is itself the real clinical decision target and was prespecified.
  • Ask whether missing follow-up was handled in a way consistent with the estimand rather than simply convenient for the table shell.

Reviewer Red Flags

  • The paper reports only change scores and never shows the adjusted follow-up estimate.
  • Percent improvement is emphasized without any discussion of low or unstable baseline denominators.
  • A responder cutoff appears clinically convenient and statistically undocumented.
  • Baseline imbalance is waved away with "the trial was randomized" even though the sample is modest and the outcome is highly baseline-dependent.
  • Missing follow-up is handled casually, with no explanation of whether the chosen model aligns with the estimand and censoring pattern.
  • The discussion talks as if a more dramatic analysis is automatically a more clinically meaningful one.

The Practical Bottom Line

Baseline adjustment in randomized trials is not glamorous, but it is one of those quiet places where methodological discipline either sharpens the answer or lets custom and convenience take the wheel.

If you are reviewing a trial, protocol, or AI-generated evidence summary that leans hard on change scores, percent improvement, or a suspiciously photogenic responder analysis, Aqrab can help stress-test the estimand, analysis choice, and red flags before the endpoint starts narrating fiction. If you want to build those critique checks into your own review pipeline, the developer tools are the cleaner place to start.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive