The Treatment Paradox: When Good Care Hides Baseline Risk
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
A high blood pressure reading warns the clinician, treatment starts, and the patient avoids the complication. In routine data, the reading may now look only weakly associated with the outcome. The predictor did not fail. Care responded to it and changed what happened next.
This is the treatment paradox in prognostic modeling: predictors influence clinical action, action influences outcomes, and the observed predictor–outcome relationship no longer represents untreated risk.

The Method in One Sentence
Draw the clinical pathway from baseline predictor to treatment and from treatment to outcome before modeling; then define whether the target is prognosis under usual care or risk under a clearly specified treatment strategy.
Concrete takeaway
Before dropping a clinically important predictor because its coefficient looks small, ask whether that predictor caused clinicians to prevent the very outcome being modeled.
Why the Association Can Reverse
Let severity at baseline increase untreated risk. Clinicians observe severity and preferentially treat the sickest patients. If treatment is effective, treated high-risk patients can have outcomes similar to—or better than—untreated lower-risk patients. A model trained on the observed outcome learns the joint behavior of biology and the care system.
That model can still predict outcomes under the same stable care pathway. But it does not automatically estimate what would happen without treatment, after treatment policy changes, or in a hospital with different access and escalation thresholds.
A Small Teaching Example
| Baseline group | Untreated event risk | Treatment uptake | Observed event risk |
|---|---|---|---|
| Lower risk | 10% | 10% | about 9.5% |
| Higher risk | 30% | 80% | about 18% |
Hypothetical assumptions: treatment halves risk; no confounding, nonadherence, competing events, or loss to follow-up. The example illustrates distortion, not an estimator.
The high-risk group still has more observed events, but the observed contrast—18% versus 9.5%—is far smaller than the untreated contrast—30% versus 10%. A sufficiently effective, selectively deployed treatment could flatten or reverse the association.
“Adjust for Treatment” Is Not a Universal Repair
Entering treatment as an ordinary covariate can be inadequate when treatment starts after baseline, changes over time, depends on evolving prognosis, and alters later measurements. Conditioning on post-baseline treatment can also change the question and introduce selection or collider bias.
Possible designs include restricting to a treatment-naive target population, using an appropriate randomized control arm, modeling risk under observed usual care, or estimating risk under specified treatment strategies with causal methods such as standardization or marginal structural models. Each targets a different estimand and requires assumptions that should be explicit.
Prediction and Counterfactual Prediction Are Different
A conventional prognostic model estimates what is likely to happen under the care patterns represented in its data. A counterfactual prediction asks what would happen under a particular action or policy. High discrimination under current practice does not validate the second claim.
Do not edit a patient's input in a risk calculator and interpret the changed output as the effect of changing that input. A strong predictor need not be causal, and a weak predictor may have been neutralized by treatment. Treatment recommendations require treatment-effect evidence, not just outcome-risk prediction.
Why Deployment Can Break the Model
Once a model is deployed, its alert may itself trigger care. Successful intervention changes the outcome labels that future monitoring observes. Calibration can drift because the model altered the pathway, not because the original algorithm was technically defective.
Evaluation therefore needs a care-process map: who saw the score, what action followed, how quickly, for whom, and whether treatment policy changed. An impact trial asks whether using the model improves decisions or outcomes; a retrospective AUC does not.
The 90-Second Treatment-Paradox Audit
- What decision will the model inform: prognosis under usual care, no treatment, or a specified strategy?
- Could any candidate predictor trigger treatment, monitoring, referral, or rescue care during follow-up?
- Were those post-baseline actions measured with their timing and dose?
- Is the outcome observed after clinicians had access to the predictors?
- Does the analysis merely adjust for treatment, or state the causal assumptions needed for that adjustment?
- Was performance evaluated in the setting and care pathway where the model will actually be used?
What Reviewers Should Demand
Require authors to define the target population, prediction time, horizon, intended decision, treatment regime, and care setting. Ask whether clinicians knew the predictors and whether those predictors influenced post-baseline intervention. Treatment timing and intensity belong in the data description, not only in a limitations paragraph.
Then match the claim to the design. “Predicts outcomes under this care pathway” is narrower than “estimates untreated risk.” “Supports treatment decisions” is stronger still and needs evidence that the decision rule improves net benefit or patient outcomes.
Sources and Evidence Maturity
- Cheong-See et al., Prediction models in obstetrics: understanding the treatment paradox and potential solutions (2016) — peer-reviewed methods article; established conceptual guidance, with solutions dependent on the target question and assumptions.
- Sperrin et al., Using marginal structural models to adjust for treatment drop-in (2018) — peer-reviewed methods study with simulation and applied illustration; supports time-varying treatment adjustment under stated causal assumptions.
- Riley et al., When and how to use data from randomised trials to develop or validate prognostic models (2019) — peer-reviewed methodological guidance; mature principles for aligning trial arms and prognosis targets.
- Efthimiou et al., Developing clinical prediction models: a step-by-step guide (3 September 2024) — peer-reviewed contemporary guidance; emphasizes the intended population, outcome, setting, users, decisions, performance, and clinical usefulness.
- Collins et al., TRIPOD+AI statement (16 April 2024) — international consensus reporting guideline; mature reporting standard, not a risk-of-bias tool or causal-identification method.
Evidence note: the numerical example is intentionally hypothetical. The causal correction needed in a real study depends on treatment timing, measured confounding, positivity, consistency, model specification, and the precise prediction target.
Where Aqrab Fits
Aqrab can help reviewers extract the prediction target, index time, treatment timing, intended use, data setting, performance measures, and stated limitations across a protocol and paper. Use that structure to expose a hidden care pathway—not to turn prediction into a causal treatment recommendation. Try Aqrab on a prediction-model paper, or explore plans for repeatable evidence-review workflows.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
AI Before–After Studies: When Faster Care Is Not Yet an AI Effect
A practical guide to evaluating healthcare AI after deployment. Audit pre-trends, concurrent comparisons, co-interventions, outcome measurement, and the claim ceiling of before–after and interrupted time-series designs.
AI Surveillance Models: Why a High AUC Cannot Justify Fewer Follow-Up Visits
A practical guide to evaluating AI surveillance models. Learn why high AUC is not enough to reduce follow-up, and audit calibration, thresholds, missed failures, utility, and prospective impact.
Prediction vs Causation: Why Your Best Risk Model Still Cannot Tell You What to Treat
A practical guide for clinical researchers on the difference between prediction and causation. Covers why strong risk models do not identify treatment effects, how to frame the right estimand, and what reviewers should flag in AI-driven clinical studies.