← Back to Blog
Causal InferenceConfoundingMethods Critique

Simpson’s Paradox: When Every Subgroup Says One Thing and the Total Says the Opposite

July 22, 2026·14 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A new therapy has a higher recovery rate than standard care among mild patients. It also has a higher recovery rate among severe patients. Every subgroup agrees: the therapy helps. Then someone pools the two groups into one table, and the therapy’s overall recovery rate comes out lower than standard care. Nobody made an arithmetic mistake. This is Simpson’s paradox, and it is one of the few pieces of statistics that can make a careful clinician feel like the ground has moved.

The reversal itself is not the hard part — it is a mechanical consequence of unequal group sizes, and you can conjure it on demand. The hard part is the question it forces: you now have two correct tables that disagree, so which one is the answer? The instinct drilled into most of us is “always look at the subgroups.” That instinct is wrong often enough to be dangerous. The table you should believe depends entirely on what the splitting variable is in the causal story — and no amount of staring at the data will tell you that.

The Core Decision Rule

Simpson’s paradox is not a statistical error to be detected and corrected. It is a fork in the road where the arithmetic goes silent and the causal question has to speak.

Decision rule:

When pooled and stratified estimates disagree, do not ask which number is bigger or which sample is larger. Ask what the stratifying variable is on the causal diagram. Adjust for it only if it is a confounder. If it is a mediator or a collider, adjusting is the mistake, and the pooled comparison is the honest one.

Why the Reversal Happens at All

A pooled rate is a weighted average, and the weights are the group sizes. If the treated patients are mostly the severe ones — who recover less often no matter what — then the treated arm carries the heavier weight on the low-recovery stratum, and its overall rate is dragged down. The control arm, stacked with mild patients, floats up. The within-group advantage of treatment is still there in both strata; it is simply outrun by the fact that the two arms are not comparable populations.

That is the whole mechanism: an unequal mix across strata that differ in baseline outcome. It is the same machinery as confounding, which is why the paradox shows up constantly in observational data, where who gets treated is never random.

See It Move

The explorer below lets you build the reversal yourself and then, crucially, decide what the stratifier means. The recovery numbers never change when you switch its role — only the verdict about which table you are allowed to report. That gap between fixed arithmetic and shifting conclusion is the entire lesson.

Interactive Simpson’s paradox explorer

The same table, one arithmetic, two opposite conclusions

Build the paradox with the sliders: a treatment that helps inside every subgroup can still look harmful once the groups are pooled. Then pick what the stratifier is in your causal story and watch which table you are allowed to believe flip — without a single number changing.

ReversalSign flip presentWithin each stratum: +8.0 ptsPooled: -32.5 pts

The honest, equal effect of treatment inside both the mild and severe groups. It is always positive here.

How much lower the severe stratum’s baseline recovery is than the mild stratum’s (80%).

How unevenly the arms sit across strata. At zero the arms are balanced and the paradox vanishes; near the top the treated are almost all severe.

Recovery rateTreatedControlDifference
Mild stratum88%80%+8.0 pts
Severe stratum38%30%+8.0 pts
Pooled (everyone)43%75%-32.5 pts

What is the stratifier in your causal story?

The numbers above do not move when you change this. Only the answer to “which table should I report?” does.

Trust the within-stratum effect. Adjust for the stratifier.

A confounder is a common cause of who gets treated and who recovers: sicker patients were both more likely to be treated and less likely to recover. That open backdoor path is exactly what the pooled comparison leaves in. Conditioning on the stratifier closes it, so the within-stratum effect is the honest one.

Answer to report: the within-stratum effect (+8.0 pts).

The point that survives

  • The reversal is real arithmetic, not a mistake — both tables are computed correctly.
  • “Always disaggregate” is wrong: adjusting for a mediator or collider is the error, not the fix.
  • Nothing in the data tells you the stratifier’s role. That is a claim you bring from a diagram.

The Same Numbers, Three Different Right Answers

Here is the part that textbooks often skip. The reversal in the table is compatible with three completely different causal structures, and each one demands a different decision. The counts are identical; the correct action inverts.

If the stratifier is a…Because…Report the…
ConfounderA common cause of who is treated and who recovers (e.g. baseline severity). It opens a backdoor path that only adjustment can close.Stratified (adjusted) effect
MediatorTreatment changes the stratifier, which then changes the outcome. Splitting on it removes the part of the effect that travels through it.Pooled (total) effect
ColliderA common effect of treatment and outcome (or their causes). It carries no confounding, but conditioning on it invents a spurious association inside each stratum.Pooled (unconditioned) effect

This is why “the more granular table is more trustworthy” is a myth. Granularity is not virtue. Conditioning on a mediator amputates a real effect; conditioning on a collider manufactures a fake one. In two of the three cases, the subgroup view is the one lying to you.

The Textbook Case, and a Modern One

Case

Kidney stones, and a dashboard that ranks two hospitals

The classic illustration is the 1986 kidney-stone comparison: open surgery beat a less invasive procedure for small stones and for large stones separately, yet lost overall, because surgeons reserved open surgery for the hard cases. Severity is a textbook confounder, so the stratified answer wins: open surgery was better.

Now the modern version. A quality dashboard compares two hospitals and adjusts for “admitted to the ICU,” a variable recorded after arrival. If ICU admission is partly a consequence of how the hospital manages a deteriorating patient, it is a mediator — and adjusting for it hides exactly the care difference you were trying to measure. Same paradox on the screen, opposite correct move. The data cannot tell the two cases apart. Only knowing when the variable was set, and by what, can.

Why Smart Analysts Fall for It

Subgroups feel safer

Disaggregating looks like diligence. “Controlling for” more variables reads as more careful, so the stratified table is trusted by reflex — even when the variable should never have entered the model.

The software is agnostic

A regression happily adjusts for a confounder, a mediator, and a collider in the same line of code. The output looks identical. Nothing flags that one of those adjustments just wrecked the estimate.

Timing gets ignored

Whether a variable is a confounder or a mediator often comes down to whether it was set before or after treatment. That timing lives in the protocol, not the dataset, and is easy to skip.

Simpson’s Paradox Is Not Non-Collapsibility

A quick but important clarification, because the two are constantly confused. Simpson’s paradox is a genuine change in the causal story: the pooled and conditional effects disagree because of how the groups are mixed, and at least one of them is answering the wrong question. Non-collapsibility is different — it is a quirk of certain effect measures, chiefly the odds ratio and hazard ratio, where the conditional and marginal estimates differ even with no confounding at all, simply because of the math of the measure.

The practical tell: a risk difference or risk ratio is collapsible, so if those reverse, you are looking at a real Simpson’s reversal driven by group mix. If only the odds ratio shifts while the risk difference holds steady, you may be seeing non-collapsibility, not confounding. Diagnosing one as the other sends you adjusting for a problem that was never there.

Reviewer Red Flags

What a defensible paper does

  • States, before looking at the data, which variables are confounders to adjust for and why.
  • Distinguishes pre-treatment covariates from post-treatment variables and refuses to adjust for the latter without a mediation model.
  • Names the estimand: total effect versus a specific conditional or direct effect.
  • Uses a collapsible measure when it wants to reason about who-changes-whom, not just the odds ratio.

What should make you nervous

  • “After adjusting for everything available” — a kitchen-sink model with no causal rationale.
  • A post-baseline variable (ICU stay, dose received, adherence) sitting in the adjustment set.
  • A subgroup reversal reported as automatically more valid than the overall result.
  • A sign flip in the odds ratio treated as confounding without checking the risk difference.

Decision Rules That Travel Well

  1. When pooled and stratified disagree, reach for the diagram before reaching for a conclusion.
  2. Ask when the stratifier was measured: before treatment points toward confounder, after points toward mediator or collider.
  3. Adjust for confounders. Do not adjust for mediators or colliders unless you are running an explicit mediation or bias analysis that accounts for them.
  4. Write down the estimand you actually want — total effect or direct effect — before you choose the table.
  5. If the reversal is only in the odds ratio, check a collapsible measure before blaming confounding.

Where Aqrab Fits

Simpson’s paradox is dangerous precisely because both tables are computed correctly, so nothing in a statistical check catches it. The failure is upstream, in whether the adjustment set matched the causal roles of the variables. Aqrab is built to read a methods section the way a careful reviewer would: flagging a post-treatment variable in the adjustment set, an unnamed estimand, or a subgroup claim leaning on a variable whose causal role was never argued.

If you want a second pass on whether your model is adjusting for the right things — and not conditioning its way into a reversal — start with Aqrab Try and make the causal role of every covariate explicit before the paradox makes the decision for you.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive