← Back to Blog
Causal InferenceReal-World EvidenceMethods Critique

Consistency and Treatment Versioning: When One Exposure Label Hides Several Different Interventions

June 12, 2026·16 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Clinical papers often write the exposure as if it were a single, clean switch: started the drug, received the pathway, underwent enhanced follow-up. But real-world interventions usually arrive in versions: different doses, different timing, different monitoring, different cointerventions, and different delivery teams.

That matters because the consistency assumption is not a philosophical decoration. It is the rule that lets an observed outcome under the intervention actually stand in for the counterfactual outcome you claim to estimate. If one label hides several materially different interventions, the causal question is underspecified before you fit a single model.

The Core Decision Rule

If you cannot describe what clinical team would need to do so that two patients both count as receiving the same intervention, your estimand is probably too vague for a confident causal interpretation.

Decision rule:

Do not treat intervention definition as solved just because treatment status is captured in the data. A measured exposure can still hide incompatible versions of care.

What Consistency Is Actually Asking You to Believe

The intervention is well-defined

The treatment label corresponds to a clinically intelligible action, not a loose bucket of loosely related practices.

Versions do not matter, or are explicitly averaged

Either the versions are irrelevant to outcome, or the estimand clearly defines the distribution of versions being averaged over.

Observed care matches the claimed contrast

If the manuscript says “initiation of treatment,” readers should not later discover the effect depends on hidden titration support, bundled monitoring, or site-specific workflow extras.

Where Versioning Sneaks In

Drugs are not single versions of themselves

Dose, route, induction schedule, refill timing, discontinuation counseling, and adverse-event monitoring can all change outcomes while still being filed under the same exposure label.

Care pathways bundle cointerventions

A pathway may include pharmacist outreach, lab reminders, and nurse triage in one hospital but only an order-set default in another.

Calendar time quietly changes the version mix

Early adopters may deliver an intervention differently from later scaled deployment, even when the EHR variable remains identical.

Transportability breaks before modeling starts

Two systems can estimate different effects simply because they are averaging over different versions of the same named treatment.

A Concrete Clinical Example

Case

“Treatment initiation” means different things across sites

Imagine a real-world study of a newly adopted cardiometabolic therapy. In Site A, initiation usually comes with protocolized titration, pharmacy outreach, side-effect counseling, and rapid follow-up. In Site B, initiation is mostly a prescription event with routine follow-up and less structured support.

Both sites may code patients as treated. But the causal contrast is not the same. If Site A reports a larger benefit, that difference is not automatically confounding, poor overlap, or model failure. It may reflect that the intervention version itself changed.

This is where many papers become overconfident. They defend exchangeability and positivity, then act as if intervention definition no longer matters. It still does.

Interactive versioning explorer

Watch the same treatment label change meaning as the intervention mix shifts

This toy example assumes the untreated outcome risk is the same across sites. What changes is the hidden mix of treatment versions under the label. One site delivers a high-touch version with stronger effect. The other mostly delivers a lighter version.

Versioning gap5.0 pointsDifference in average risk reduction across sites

If two papers both say they studied “treatment initiation,” but the version mix differs this much, the causal question is not as portable as the shared label suggests.

Site A average treated risk

21%

Average absolute risk reduction: 11.5 points

Site B average treated risk

26%

Average absolute risk reduction: 6.5 points

QuantityValueWhy it matters
High-touch version treated risk18%This is the risk if everyone received the stronger, better-supported intervention version.
Low-touch version treated risk28%This is the risk if everyone received the lighter version hiding under the same exposure label.
Site A pooled effect11.5 pointsAn average effect is only interpretable if the version mix is part of the intervention definition.
Site B pooled effect6.5 pointsA second site can produce a different answer even with equal baseline risk and no confounding change.
Versioning gap5.0 pointsThis is ambiguity from intervention definition, not automatically from bias correction failure.

When a Vague Exposure Is Still Defensible

Exposure labelCan it still work?What must be explicit
A specific dose and start windowUsually yesHow initiation, adherence support, and allowed cointerventions were operationalized.
Any prescription recordedOnly with cautionWhy heterogeneous versions can reasonably be averaged and what version mix the result represents.
A multi-step care pathwaySometimesWhich elements are mandatory, optional, site-specific, or time-varying.
A policy label that drifted over calendar timeRarely without redesignHow the intervention changed over time and whether the estimand is era-specific rather than pooled.

Reviewer Red Flags

Signals the intervention is underspecified

  • The exposure is defined by a billing code or prescription record with little clinical description.
  • The study pools sites or years where the delivery workflow clearly changed.
  • Dose, timing, route, monitoring, or cointerventions are left vague because they are “implementation details.”
  • The paper interprets a pooled effect as portable to new settings without describing the version mix.
  • “Real world” is used as a license to stop defining the intervention precisely.

Better reviewer questions

  • What exactly had to happen for a patient to count as treated?
  • Which treatment versions were possible, and were they clinically interchangeable for the outcome?
  • Did the version mix differ by site, calendar time, prescriber type, or severity level?
  • If the effect is an average over versions, what population and workflow does that average represent?
  • Would a clinician know how to reproduce the intervention from the paper alone?

Why This Matters for Aqrab

Aqrab is most useful before a team gets emotionally attached to a vague exposure definition. The hard work is not only running the model. It is forcing the protocol to say what intervention is actually being contrasted, what hidden versions exist, and whether the effect is meant to travel outside the workflow that produced it.

If your group wants a faster way to interrogate intervention definitions, estimands, and reviewer failure modes before manuscript drafting, Aqrab's study-design critique workflow is built for exactly that pressure test.

Bottom Line

Consistency does not ask for a perfect world. It asks you to name the intervention honestly. If one exposure label hides several materially different versions of care, the causal effect is only as clear as the version definition you are willing to write down.