← Back to Blog
Study DesignBias DiagnosticsMethods Critique

Regression to the Mean: When an Extreme Baseline Improves by Itself

September 24, 2026·12 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A clinic enrolls patients only when their blood pressure, pain score, or service use is unusually high. The same outcome is measured after an intervention, and the average is lower. The tempting conclusion is that the intervention worked. But an extreme observation often contains a temporary fluctuation or measurement error that is unlikely to repeat.

Regression to the mean is that statistical tendency: when repeated measurements are imperfectly correlated, units selected for an extreme first value tend to have a less extreme second value on average—even without an effective intervention.

The Method in One Sentence

When enrollment depends on an extreme baseline value, estimate the intervention effect against a concurrent group selected and measured the same way; do not interpret the treated group's pre-post change by itself.

Concrete takeaway

A paired test can show that patients changed. It cannot show how much of that change the intervention caused. The causal contrast is the change beyond what comparable patients would have shown without the intervention.

Why the Extreme Value Softens

Think of a measurement as a stable component plus short-term fluctuation and measurement error. Selection on a high observed value preferentially captures people whose temporary component happened to point upward. At the next visit, that temporary component may be smaller or point the other way, so the group average moves toward its longer-run mean.

The effect becomes more visible when measurement reliability is lower, natural variability is larger, and eligibility uses a more extreme threshold. It is not proof that every individual improves, and it is not a biological law that outcomes must recover. It is a predictable consequence of selecting on a noisy extreme.

A Small Teaching Example

GroupMean baselineMean follow-upMean change
Selected intervention group160150−10
Comparable untreated group159150−9
Between-group contrast in change——−1

Hypothetical values for a continuous outcome where lower is better. The example isolates interpretation; it is not a clinical effect estimate or a recommended analysis.

The intervention group improved by 10 units, and a paired test might label that change statistically significant. Yet the untreated comparison improved by 9 units under the same extreme-value selection. The comparative signal is only 1 unit. Without the comparator, the apparent 10-unit benefit mixes any treatment effect with regression to the mean, natural history, co-interventions, and measurement changes.

A Paired Test Does Not Repair the Design

A paired analysis accounts for the correlation between two measurements from the same person. It asks whether the mean within-person change differs from zero under its assumptions. Zero change is not the counterfactual outcome that would have occurred without treatment.

If the selected group would have improved anyway, a precise paired estimate can be precisely wrong as an intervention effect. Statistical significance does not separate treatment from regression to the mean.

Better Design Before Better Modeling

Random allocation after the same eligibility assessment allows regression to the mean to operate in both arms; the randomized between-arm contrast protects the treatment estimate. When randomization is infeasible, a credible concurrent comparator should face the same selection rule, calendar time, measurement process, and follow-up.

Repeated baseline measurements can reduce dependence on one extreme reading and help characterize variability. An interrupted time-series design with enough observations before and after an intervention can distinguish an intervention-related level or slope change from an already moving trajectory, though it still needs assumptions about concurrent events and outcome measurement.

Analysis Must Match the Design

In a comparative study, estimate the between-group treatment contrast and account appropriately for baseline outcome values. In randomized trials with continuous outcomes, baseline-adjusted analysis is often more efficient than testing change separately within each arm. The important protection, however, comes from the comparison created by design—not from a magical regression term.

When only one treated group has one pre-intervention and one post-intervention measurement, no statistical adjustment can reconstruct an unobserved no-treatment trajectory without strong, untestable assumptions. Label the finding as pre-post change and keep the causal claim narrow.

Do Not Confuse It With Related Problems

Natural history is real change in the underlying condition. Measurement error is one contributor to imperfect repeatability. Regression to the mean is the average pattern created when selection uses an extreme noisy value. Confounding is a different source of noncomparability, although all can coexist in an uncontrolled study.

A comparison group helps reveal the combined no-treatment change, but a poor comparator can create new bias. Review who entered each group, why, when, and how outcomes were measured.

The 90-Second Regression-to-the-Mean Audit

  • Were participants enrolled because their baseline outcome was unusually high or low?
  • Was eligibility based on one noisy measurement or on repeated confirmation?
  • Is there a concurrent comparison group selected and measured in the same way?
  • Does the paper treat a within-group change or paired-test p-value as an intervention effect?
  • Could symptoms, laboratory values, or service use fluctuate naturally over time?
  • Were outcome measurement, follow-up timing, co-interventions, and attrition comparable between groups?

What Reviewers Should Demand

Ask for the eligibility threshold, number and timing of baseline measurements, test-retest reliability where known, participant flow around the threshold, concurrent comparator, and a direct between-group effect estimate with uncertainty.

Reject language that turns “the treated group improved” into “the treatment caused improvement” without a defensible counterfactual. The more strongly entry depends on an extreme value, the more prominently regression to the mean belongs in the design and interpretation.

Sources and Evidence Maturity

Evidence note: the table is intentionally hypothetical. The magnitude of regression to the mean in a real dataset depends on selection, repeatability, outcome distribution, timing, and the target population; it cannot be inferred from the example.

Where Aqrab Fits

Aqrab can help reviewers extract the eligibility threshold, baseline and follow-up timing, comparator, outcome measurement, analysis population, and reported effect contrast across a protocol and paper. Use that structure to expose an uncontrolled change claim—not to manufacture a missing counterfactual. Try Aqrab on a pre-post study, or explore plans for repeatable evidence-review workflows.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive
Next guide

This is the newest guide so far.