← Back to Blog
Causal InferenceBias DiagnosticsMethods Critique

M-Bias: When “Adjusted for Severity” Is Actually the Problem

June 28, 2026·15 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Clinical researchers are trained to feel virtuous about adjustment. If a study controlled for a severity score, referral marker, or utilization variable, the model can look more responsible at first glance. Sometimes it is. Sometimes it is the moment the estimate stops meaning what the abstract says it means.

M-bias is the warning that not every pre-exposure covariate is a confounder. A variable can sit at the meeting point of two upstream causes, and conditioning on it can open a non-causal path between treatment and outcome. The result is a manufactured association wearing the costume of careful adjustment.

The Core Decision Rule

Do not include a baseline variable just because it feels clinically important or predicts cohort entry. Include it only if you can explain how it blocks a real backdoor path rather than opens one.

Decision rule:

If a variable is a common effect of two upstream processes, adjusting for it can create bias even when treatment truly has no effect on the outcome.

What Makes It “M” Bias

The label comes from the rough shape of the causal graph. One latent cause pushes treatment and the adjustment variable. Another latent cause pushes the outcome and the same adjustment variable. Those two latent causes may be independent until you condition on their shared consequence.

Toy graph

Severity preference → Treatment
Severity preference → Recorded severity score ← Frailty burden
Frailty burden → Outcome

Before conditioning, treatment and outcome do not become associated through that path. Condition on the recorded severity score, and the path opens.

Why Smart Analysts Fall for It

The variable sounds clinical

Severity scores, referral urgency, test positivity, and care intensity feel too important to omit.

It is measured before treatment

Baseline timing helps, but baseline alone does not prove a variable is a confounder.

The model fit improves

Better prediction, tighter standard errors, or stronger balance diagnostics can coexist with worse causal identification.

A Concrete Clinical Example

Case

Antiarrhythmic treatment in a tertiary referral arrhythmia clinic

Imagine an observational comparison of two antiarrhythmic strategies. One latent process is clinician preference for aggressive rhythm control. Another is underlying frailty and arrhythmia complexity. Both influence whether a patient reaches the tertiary clinic and how severe they look in the intake score.

If the analysis adjusts for that intake severity score without clarifying what upstream processes built it, the study may open a path between treatment choice and prognosis that was not there in the source population. The estimate can then look cleaner, narrower, and more wrong.

Interactive M-bias explorer

Watch a null effect become publishable after collider selection

In this toy world, treatment does not cause the outcome. Two separate latent causes drive treatment and outcome, and both also influence who enters the analyzed sample. Restrict the cohort hard enough and an association appears out of structure alone.

Current readThe selected sample is manufacturing a visible storyFull cohort RD: -0.2 ppSelected sample RD: -6.6 pp

Higher values mean the selection process depends more heavily on both latent causes.

Lower values mimic tighter restriction to a referred, hospitalized, tested, or complete-case subgroup.

More noise weakens the selection mechanism. Less noise makes the collider structure easier to see.

SamplePatientsOutcome risk if treatedOutcome risk if untreatedApparent effect
Full cohort4,50050.4%50.6%-0.2 pp
Collider-selected sample1,57667.2%73.8%-6.6 pp

What to notice

The apparent treatment-outcome contrast is no longer coming from a causal effect. It is being induced by conditioning on who entered the analysis.

The full cohort stays near the truth because treatment and outcome were generated independently. The selected sample drifts because entry depends on both upstream causes.

Clinical translation

  • Hospitalized-only cohorts can do this.
  • Tested-only cohorts can do this.
  • Complete-case analyses can do this when missingness depends on prognosis and care.
  • Adding more regression terms after selection does not restore the source population.

M-Bias vs Confounding vs Collider Bias

ProblemWhat the variable is doingWhat the analyst should do
ConfoundingThe variable is a common cause of treatment and outcome.Block that backdoor path with design or justified adjustment.
Collider biasThe variable is a common effect of treatment and outcome.Do not condition on it unless the estimand explicitly requires a selected population.
M-biasThe variable is a common effect of separate upstream causes of treatment and outcome.Treat the variable as a collider candidate, not an automatic severity adjustment.

Reviewer Red Flags

What a defensible paper does

  • Explains cohort entry, referral, testing, and severity measurement as causal processes.
  • Shows a DAG or equivalent subject-matter logic for why each adjusted variable belongs.
  • Separates confounders from selection variables and proxies.
  • Uses design restraint when selected cohorts make collider structures plausible.

What should make you nervous

  • The paper adjusts for a severity score without explaining how the score was generated.
  • The analysis is restricted to a tested, referred, hospitalized, or complete-case subgroup.
  • Balance or c-statistics are used as proof that the covariate set is causally correct.
  • The discussion treats “baseline covariate” as synonymous with “safe to adjust for.”

Decision Rules That Travel Well

  1. Ask what caused the variable, not just when it was measured.
  2. Map cohort entry and measurement processes before fitting the model.
  3. Be suspicious of variables that summarize referral, care intensity, or selection into the dataset.
  4. Use prediction performance as a secondary diagnostic, not a causal passport.
  5. If you cannot justify the adjustment set in words, the software cannot justify it for you.

Where Aqrab Fits

M-bias is the kind of failure mode that hides inside polished methods sections. Aqrab is useful when you need a second pass on whether a variable is solving confounding, encoding selection, or quietly opening a path you did not intend to study.

If you want a structured critique of adjustment logic before submission, start with Aqrab Try and make the causal diagram explicit before the regression output becomes the story.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive