← Back to Blog
Prediction ModelsClinical AIStudy Design

The Treatment Paradox: When Good Care Hides Baseline Risk

September 24, 2026·12 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A high blood pressure reading warns the clinician, treatment starts, and the patient avoids the complication. In routine data, the reading may now look only weakly associated with the outcome. The predictor did not fail. Care responded to it and changed what happened next.

This is the treatment paradox in prognostic modeling: predictors influence clinical action, action influences outcomes, and the observed predictor–outcome relationship no longer represents untreated risk.

A high-risk gauge leads to a treatment shield and then a lower adverse-outcome curve, while the naive direct path from baseline risk to observed outcome is crossed out
Observed outcomes contain the effects of care delivered after the predictor was seen. A model must say whether it predicts usual-care risk or risk under a specified intervention strategy.

The Method in One Sentence

Draw the clinical pathway from baseline predictor to treatment and from treatment to outcome before modeling; then define whether the target is prognosis under usual care or risk under a clearly specified treatment strategy.

Concrete takeaway

Before dropping a clinically important predictor because its coefficient looks small, ask whether that predictor caused clinicians to prevent the very outcome being modeled.

Why the Association Can Reverse

Let severity at baseline increase untreated risk. Clinicians observe severity and preferentially treat the sickest patients. If treatment is effective, treated high-risk patients can have outcomes similar to—or better than—untreated lower-risk patients. A model trained on the observed outcome learns the joint behavior of biology and the care system.

That model can still predict outcomes under the same stable care pathway. But it does not automatically estimate what would happen without treatment, after treatment policy changes, or in a hospital with different access and escalation thresholds.

A Small Teaching Example

Baseline groupUntreated event riskTreatment uptakeObserved event risk
Lower risk10%10%about 9.5%
Higher risk30%80%about 18%

Hypothetical assumptions: treatment halves risk; no confounding, nonadherence, competing events, or loss to follow-up. The example illustrates distortion, not an estimator.

The high-risk group still has more observed events, but the observed contrast—18% versus 9.5%—is far smaller than the untreated contrast—30% versus 10%. A sufficiently effective, selectively deployed treatment could flatten or reverse the association.

“Adjust for Treatment” Is Not a Universal Repair

Entering treatment as an ordinary covariate can be inadequate when treatment starts after baseline, changes over time, depends on evolving prognosis, and alters later measurements. Conditioning on post-baseline treatment can also change the question and introduce selection or collider bias.

Possible designs include restricting to a treatment-naive target population, using an appropriate randomized control arm, modeling risk under observed usual care, or estimating risk under specified treatment strategies with causal methods such as standardization or marginal structural models. Each targets a different estimand and requires assumptions that should be explicit.

Prediction and Counterfactual Prediction Are Different

A conventional prognostic model estimates what is likely to happen under the care patterns represented in its data. A counterfactual prediction asks what would happen under a particular action or policy. High discrimination under current practice does not validate the second claim.

Do not edit a patient's input in a risk calculator and interpret the changed output as the effect of changing that input. A strong predictor need not be causal, and a weak predictor may have been neutralized by treatment. Treatment recommendations require treatment-effect evidence, not just outcome-risk prediction.

Why Deployment Can Break the Model

Once a model is deployed, its alert may itself trigger care. Successful intervention changes the outcome labels that future monitoring observes. Calibration can drift because the model altered the pathway, not because the original algorithm was technically defective.

Evaluation therefore needs a care-process map: who saw the score, what action followed, how quickly, for whom, and whether treatment policy changed. An impact trial asks whether using the model improves decisions or outcomes; a retrospective AUC does not.

The 90-Second Treatment-Paradox Audit

  • What decision will the model inform: prognosis under usual care, no treatment, or a specified strategy?
  • Could any candidate predictor trigger treatment, monitoring, referral, or rescue care during follow-up?
  • Were those post-baseline actions measured with their timing and dose?
  • Is the outcome observed after clinicians had access to the predictors?
  • Does the analysis merely adjust for treatment, or state the causal assumptions needed for that adjustment?
  • Was performance evaluated in the setting and care pathway where the model will actually be used?

What Reviewers Should Demand

Require authors to define the target population, prediction time, horizon, intended decision, treatment regime, and care setting. Ask whether clinicians knew the predictors and whether those predictors influenced post-baseline intervention. Treatment timing and intensity belong in the data description, not only in a limitations paragraph.

Then match the claim to the design. “Predicts outcomes under this care pathway” is narrower than “estimates untreated risk.” “Supports treatment decisions” is stronger still and needs evidence that the decision rule improves net benefit or patient outcomes.

Sources and Evidence Maturity

Evidence note: the numerical example is intentionally hypothetical. The causal correction needed in a real study depends on treatment timing, measured confounding, positivity, consistency, model specification, and the precise prediction target.

Where Aqrab Fits

Aqrab can help reviewers extract the prediction target, index time, treatment timing, intended use, data setting, performance measures, and stated limitations across a protocol and paper. Use that structure to expose a hidden care pathway—not to turn prediction into a causal treatment recommendation. Try Aqrab on a prediction-model paper, or explore plans for repeatable evidence-review workflows.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive