← Back to Blog
Bias DiagnosticsStudy DesignMethods Critique

The Will Rogers Phenomenon: When Better Staging Improves Every Group and Nobody Lives Longer

July 27, 2026·13 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

The comedian Will Rogers is supposed to have said that when the Okies left Oklahoma and moved to California during the Dust Bowl, they raised the average intelligence of both states. It is a joke about arithmetic: move the least-clever Oklahomans across a state line where they become the least-clever Californians, and the mean rises on both sides — without a single person getting any smarter. In 1985, Alvan Feinstein and colleagues borrowed the line to name a trap that had been quietly inflating cancer survival statistics for decades.

Here is the clinical version. A hospital adopts better imaging. Suddenly some patients who used to be classified as early-stage — because their small metastases were invisible — are correctly reclassified as advanced-stage. Compare the new era to the old, stage by stage, and something remarkable appears: survival has improved for early-stage patients and for advanced-stage patients. Every subgroup looks better. It is tempting to credit the new treatments. But the overall survival of the whole population may not have changed at all. Nobody lived longer. The patients simply moved between boxes, and the boxes both got a better-looking average. This is the Will Rogers phenomenon, and it is one of the most seductive artifacts in clinical research because it wears the exact costume of progress.

The Core Decision Rule

Whenever you compare stage-specific, grade-specific, or risk-group-specific outcomes across two cohorts — two eras, two hospitals, two registries — your first question is not “did the subgroups improve?” It is “did the definition of the subgroups change between the two cohorts?”

Decision rule:

If a classification system reassigns patients between groups — because the test got more sensitive, or the criteria changed — then improvement within every group is consistent with no one benefiting at all. Trust a stage-specific survival gain only when the overall, unstratified outcome moved in the same direction and the stage distribution did not shift. If overall survival is flat while every stage improves, suspect migration before you suspect medicine.

Why Reclassification Lifts Both Groups at Once

The mechanism is precise and worth stating carefully, because the surprising part is that both groups improve, not just one. Consider the patients who get reclassified — the ones a sharper test moves from “early” to “advanced.” By definition they harbored occult advanced disease. That makes them two things at once:

  • The worst of the early group. Among patients called early-stage, these were the ones who were actually further along and destined to do poorly. Remove them and the early group loses its worst outcomes — so the early-stage average survival rises.
  • The best of the advanced group. Among patients who are genuinely advanced, these newly-detected ones have the least-advanced disease — smaller burden, caught sooner. Add them and the advanced group gains its best outcomes — so the advanced-stage average survival also rises.

One migration, two averages up. The only quantity that cannot be fooled is the outcome computed over everyone, ignoring the labels: it is the same set of survival times regardless of how you sort them, so it does not move. That invariant — the whole-cohort result — is your anchor. Stage-specific numbers are a story about how patients were sorted; the overall number is a story about how patients fared.

See It Move

The explorer below holds a cohort of 100 patients whose survival never changes. Old imaging leaves 20 truly-advanced patients miscounted as early. Slide the staging accuracy up and watch those patients migrate to where they belong. Both stage averages climb; the cohort-wide average stays frozen at the same value. Nothing about the patients changed — only the sorting.

Interactive stage-migration explorer

Better staging, longer survival in every group, nobody saved

The same 100 patients, with the same fixed survival, staged two ways. Old imaging cannot see occult advanced disease, so 20 patients who are truly advanced sit in the early group. As you sharpen the staging test, those patients migrate to where they belong — and both stage averages rise at once, while the cohort-wide average never moves.

Fixed truth100 patientsCohort survival: 44.4 monever changes
Custom

How many of the 20 occult-advanced patients the newer test correctly moves out of the early group and into the advanced group. Nobody is cured; they are only reclassified.

Early stage

80 patients

52.5 mo0.0 mo vs old staging

Advanced stage

20 patients

12.0 mo0.0 mo vs old staging

Whole cohort

all 100 patients

44.4 mo+0.0 mo — unchanged
Early-stage average survival52.5 mo
Advanced-stage average survival12.0 mo
Whole-cohort average survival (fixed)44.4 mo

Advanced-stage patients now make up 20% of the cohort, up from 20%— the stage distribution shifted even though the patients did not.

What just happened

This is the old-staging picture: the occult-advanced patients are still miscounted as early, dragging the early-stage average down.

  • The migrating patients were the worst outcomes among the “early” group, so removing them lifts the early average. They are the best outcomes among the “advanced” group, so adding them lifts the advanced average. One move, both means up.
  • The cohort-wide average is arithmetically pinned: it is the same 100 survival times no matter how you sort them into boxes. If overall survival is flat while every stage improves, you are looking at reclassification, not benefit.
  • This is why comparing stage-specific survival across eras or hospitals is treacherous whenever the staging technology or criteria changed in between.

Where the Problem Hides in Real Studies

Stage migration is a cancer-staging story by origin, but the phenomenon appears anywhere a graded classification gets sharper or is redefined between the groups you compare.

Where it shows upWhat changed the labelsThe misleading appearance
Cancer stage comparisons across erasPET/CT and MRI detect nodal or distant disease that older CT missed.Survival improves in every TNM stage; people conclude therapy advanced, though overall survival may be flat.
High-sensitivity troponin and MIA more sensitive assay reclassifies small events from “no MI” to “MI.”Both the MI group and the non-MI group post lower event rates, so both look lower-risk than before.
Sepsis and AKI redefinitionsNew consensus criteria move borderline patients in or out of the diagnosis.Case-fatality “improves” after a definition change with no change in care.
Biomarker-defined risk strataA new marker splits an old risk group; the reclassified patients sit at the boundary.Each new risk tier looks well-separated and prognostic, partly by construction.
“Our outcomes improved” auditsA center upgrades diagnostics, then compares this year’s stage-specific results to last year’s.Every severity band improves; leadership credits a new pathway rather than sharper labeling.

The common thread is a moving definition. When the boundary between groups is drawn differently in the two cohorts you compare, the groups are not the same groups, and their averages are not comparable — no matter how carefully everything else was measured.

How to Tell Migration From Real Progress

The phenomenon is not a reason to distrust every stage-specific improvement — treatments do get better. It is a reason to demand three specific checks before believing one.

  • Check the whole-cohort outcome. Compute survival (or the event rate) over the entire population, ignoring stage. Real therapeutic progress moves this number; pure stage migration cannot. An improvement confined to the strata, absent in the total, is the signature of migration.
  • Check the stage distribution. If a larger fraction of the cohort is now classified advanced (or a smaller fraction early) than before, patients moved between boxes. A shifting distribution across cohorts is the fingerprint you are looking for.
  • Check whether the classification changed. Did the imaging, the assay, the biopsy protocol, or the diagnostic criteria differ between the two cohorts? If yes, stage-specific comparison is off the table until you account for it.

When the classification really did change and you still need a comparison, the defensible moves are to report the unstratified outcome as the primary contrast, to stage both cohorts with a single common standard where records allow, or to restrict to a period with stable classification. What you cannot do is line up old stage-specific curves against new ones and read the gap as benefit.

How This Differs From Its Neighbors

Better diagnostic technology spawns a whole family of survival illusions, and they are routinely conflated. They are distinct mechanisms with distinct fixes.

  • Lead-time bias inflates survival measured from diagnosis: detecting the same disease earlier starts the clock sooner, so survival looks longer even if death arrives on schedule. It is about when the clock starts, not how patients are grouped.
  • Overdiagnosis and length-time bias pull indolent, slow-growing cases into the detected pool, so the detected group looks like it does unusually well. It is about which extra cases you find.
  • The Will Rogers phenomenon is about relabeling existing patients: no new cases, no clock change — the same people are re-sorted across a boundary, lifting the average on both sides. It is the only one of the three that improves every group at once while leaving the total untouched.

It is also a close cousin of Simpson’s paradox. Both are stories about how sorting patients into groups can make the parts disagree with the whole. The difference is that Simpson’s paradox compares one grouping against the pooled data at a single moment; the Will Rogers phenomenon compares the same grouping scheme applied with two different sensitivities across cohorts.

Reviewer Red Flags

What a defensible study does

  • Reports the overall, unstratified outcome alongside any stage-specific comparison.
  • States whether staging technology or criteria were identical across the compared cohorts.
  • Shows the stage distribution in each cohort so a reader can see if patients migrated.
  • Re-stages historical cases to a common standard, or restricts to a stable-classification window.

What should make you nervous

  • Survival improves in every stage, but overall survival is flat or barely mentioned.
  • The proportion classified advanced changed between eras, unremarked.
  • A new imaging modality or assay was adopted between the cohorts being compared.
  • A definition change (sepsis, AKI, MI) coincides with the reported outcome improvement.

Decision Rules That Travel Well

  1. Before believing a stage-specific improvement, look for the whole-cohort outcome; migration cannot move it.
  2. Compare the stage distributions across cohorts — a shift is direct evidence of reclassification.
  3. Ask whether the classification test or criteria changed between the groups; if so, stage-specific comparison is suspended.
  4. When definitions changed, re-stage to a common standard or make the unstratified result the headline.
  5. Keep it separate from lead-time bias and overdiagnosis — same trigger, different arithmetic, different fix.

Where Aqrab Fits

The Will Rogers phenomenon is dangerous precisely because it survives a well-run study: the measurements can be flawless and the conclusion still wrong, because the two cohorts were never labeled by the same yardstick. Catching it means reading a comparison for what it does not say — whether the staging changed between eras, whether the stage distribution drifted, and whether the overall outcome moved or only the strata did. Aqrab is built to read study design the way a careful methodologist would, surfacing when a subgroup-by-subgroup improvement rests on a classification that shifted underneath it.

If you are appraising a study that reports survival gains stage by stage across two eras — or writing one — start with Aqrab Try and check the overall outcome and the stage distribution before you let the per-stage curves tell you a story about progress.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive