Consistency and Treatment Versioning: When One Exposure Label Hides Several Different Interventions
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Clinical papers often write the exposure as if it were a single, clean switch: started the drug, received the pathway, underwent enhanced follow-up. But real-world interventions usually arrive in versions: different doses, different timing, different monitoring, different cointerventions, and different delivery teams.
That matters because the consistency assumption is not a philosophical decoration. It is the rule that lets an observed outcome under the intervention actually stand in for the counterfactual outcome you claim to estimate. If one label hides several materially different interventions, the causal question is underspecified before you fit a single model.
The Core Decision Rule
If you cannot describe what clinical team would need to do so that two patients both count as receiving the same intervention, your estimand is probably too vague for a confident causal interpretation.
Decision rule:
Do not treat intervention definition as solved just because treatment status is captured in the data. A measured exposure can still hide incompatible versions of care.
What Consistency Is Actually Asking You to Believe
The intervention is well-defined
The treatment label corresponds to a clinically intelligible action, not a loose bucket of loosely related practices.
Versions do not matter, or are explicitly averaged
Either the versions are irrelevant to outcome, or the estimand clearly defines the distribution of versions being averaged over.
Observed care matches the claimed contrast
If the manuscript says “initiation of treatment,” readers should not later discover the effect depends on hidden titration support, bundled monitoring, or site-specific workflow extras.
Where Versioning Sneaks In
Drugs are not single versions of themselves
Dose, route, induction schedule, refill timing, discontinuation counseling, and adverse-event monitoring can all change outcomes while still being filed under the same exposure label.
Care pathways bundle cointerventions
A pathway may include pharmacist outreach, lab reminders, and nurse triage in one hospital but only an order-set default in another.
Calendar time quietly changes the version mix
Early adopters may deliver an intervention differently from later scaled deployment, even when the EHR variable remains identical.
Transportability breaks before modeling starts
Two systems can estimate different effects simply because they are averaging over different versions of the same named treatment.
A Concrete Clinical Example
Case
“Treatment initiation” means different things across sites
Imagine a real-world study of a newly adopted cardiometabolic therapy. In Site A, initiation usually comes with protocolized titration, pharmacy outreach, side-effect counseling, and rapid follow-up. In Site B, initiation is mostly a prescription event with routine follow-up and less structured support.
Both sites may code patients as treated. But the causal contrast is not the same. If Site A reports a larger benefit, that difference is not automatically confounding, poor overlap, or model failure. It may reflect that the intervention version itself changed.
This is where many papers become overconfident. They defend exchangeability and positivity, then act as if intervention definition no longer matters. It still does.
Interactive versioning explorer
Watch the same treatment label change meaning as the intervention mix shifts
This toy example assumes the untreated outcome risk is the same across sites. What changes is the hidden mix of treatment versions under the label. One site delivers a high-touch version with stronger effect. The other mostly delivers a lighter version.
If two papers both say they studied “treatment initiation,” but the version mix differs this much, the causal question is not as portable as the shared label suggests.
Site A average treated risk
21%
Average absolute risk reduction: 11.5 points
Site B average treated risk
26%
Average absolute risk reduction: 6.5 points
| Quantity | Value | Why it matters |
|---|---|---|
| High-touch version treated risk | 18% | This is the risk if everyone received the stronger, better-supported intervention version. |
| Low-touch version treated risk | 28% | This is the risk if everyone received the lighter version hiding under the same exposure label. |
| Site A pooled effect | 11.5 points | An average effect is only interpretable if the version mix is part of the intervention definition. |
| Site B pooled effect | 6.5 points | A second site can produce a different answer even with equal baseline risk and no confounding change. |
| Versioning gap | 5.0 points | This is ambiguity from intervention definition, not automatically from bias correction failure. |
When a Vague Exposure Is Still Defensible
| Exposure label | Can it still work? | What must be explicit |
|---|---|---|
| A specific dose and start window | Usually yes | How initiation, adherence support, and allowed cointerventions were operationalized. |
| Any prescription recorded | Only with caution | Why heterogeneous versions can reasonably be averaged and what version mix the result represents. |
| A multi-step care pathway | Sometimes | Which elements are mandatory, optional, site-specific, or time-varying. |
| A policy label that drifted over calendar time | Rarely without redesign | How the intervention changed over time and whether the estimand is era-specific rather than pooled. |
Reviewer Red Flags
Signals the intervention is underspecified
- The exposure is defined by a billing code or prescription record with little clinical description.
- The study pools sites or years where the delivery workflow clearly changed.
- Dose, timing, route, monitoring, or cointerventions are left vague because they are “implementation details.”
- The paper interprets a pooled effect as portable to new settings without describing the version mix.
- “Real world” is used as a license to stop defining the intervention precisely.
Better reviewer questions
- What exactly had to happen for a patient to count as treated?
- Which treatment versions were possible, and were they clinically interchangeable for the outcome?
- Did the version mix differ by site, calendar time, prescriber type, or severity level?
- If the effect is an average over versions, what population and workflow does that average represent?
- Would a clinician know how to reproduce the intervention from the paper alone?
Why This Matters for Aqrab
Aqrab is most useful before a team gets emotionally attached to a vague exposure definition. The hard work is not only running the model. It is forcing the protocol to say what intervention is actually being contrasted, what hidden versions exist, and whether the effect is meant to travel outside the workflow that produced it.
If your group wants a faster way to interrogate intervention definitions, estimands, and reviewer failure modes before manuscript drafting, Aqrab's study-design critique workflow is built for exactly that pressure test.
Bottom Line
Consistency does not ask for a perfect world. It asks you to name the intervention honestly. If one exposure label hides several materially different versions of care, the causal effect is only as clear as the version definition you are willing to write down.