AI Before–After Studies: When Faster Care Is Not Yet an AI Effect
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
AI before–after studies are becoming the default real-world evaluation of clinical tools: measure care before go-live, measure it again after go-live, and credit the difference to AI. The problem is that a deployment date divides a timeline. It does not randomize it. Staffing, demand, protocols, case mix, documentation, and ordinary secular change cross that date too.
That does not make deployment evidence useless. It means the design must build a credible version of the outcome trajectory that would have occurred without deployment. The useful question is not “Did the metric improve?” It is “What else could have drawn the same curve?”
The Clean Metaphor: A Hinge Is Not a Coin Flip
A go-live date is a hinge: it separates before from after. Randomization is a coin flip: it creates comparable alternatives.
A hinge can reveal a break in the timeline. It cannot, by itself, tell you which force moved the door. That attribution comes from pre-trends, concurrent comparisons, stable measurement, and an explicit account of everything else that changed.
A 2026 Fracture-Detection Study Shows the Design at Its Best
Graif and colleagues evaluated an AI-assisted radiographic fracture-detection tool in an ambulatory orthopedic emergency department. Their retrospective study included 8,253 visits coded as fracture-excluded and 6,864 coded as fracture-confirmed from January 2021 through February 2026. The tool went live on October 1, 2023.
Mean length of stay in fracture-excluded visits fell by about eight minutes after deployment, while the fracture-confirmed internal comparison changed little. A complementary interrupted time-series analysis estimated a larger immediate level decrease against a flat pre-deployment trend. The reduction was concentrated in slower visits rather than spread evenly across the distribution, and early revisit rates did not increase.
Why this is more informative than a basic pre–post comparison
The internal comparison tests whether the change was specific to the pathway where a negative image could shorten disposition. The time series tests whether the apparent improvement was already under way. The authors still conclude that deployment was associated with shorter stays and explicitly stop short of causation. That claim boundary is a methodological strength.
Interactive design auditor
How Much Can This AI Deployment Study Claim?
Describe the evaluation, then read the claim ceiling. This is a teaching aid, not a risk-of-bias score: one unchecked assumption can matter more than several checked boxes.
Design you described
Controlled interrupted time series
Claim ceiling
The deployment series changed relative to both its prior trajectory and a concurrent comparison series. A causal interpretation still depends on comparable trends, stable measurement, and no differential co-intervention.
Still unresolved
- • An internal comparison may share the same AI exposure and workflow changes; it tests specificity, not complete absence of exposure.
- • A coding or ascertainment change at go-live can move the endpoint without changing care.
- • Staffing, triage, protocols, demand, and other software changes remain bundled with the AI launch.
Best next upgrade
- • Explain why the tool should affect one pathway but not the internal comparison, then test that mechanism.
- • Re-adjudicate a sample with the same outcome definition in both periods or validate the code-based endpoint.
- • Build a dated co-intervention log and model or bound the changes most likely to affect the outcome.
Four Designs That Should Not Share One Label
| Design | What it adds | What still carries the claim |
|---|---|---|
| Uncontrolled before–after | A temporal contrast | The implausible hope that nothing else changed |
| Controlled before–after | A concurrent change for comparison | Comparable counterfactual change across groups |
| Interrupted time series | Pre-trend, level, slope, seasonality | No concurrent shock at the interruption |
| Controlled interrupted time series | Trend plus concurrent comparison | Comparable trends and no differential co-intervention |
More observations do not automatically create stronger evidence. An interrupted time series needs a defensible interruption, enough pre-deployment data to characterize the counterfactual trend, a prespecified impact shape, and models that address seasonality and autocorrelation. A control series helps only when its relationship to deployment and to the outcome is clinically argued rather than selected because it happened to look flat.
The Internal Control Is a Mechanism Test, Not a Parallel Universe
In the fracture study, fracture-confirmed visits were not a separate hospital untouched by deployment. They experienced the same calendar time and the same tool, but the hypothesized workflow benefit was different: a rapid negative interpretation may facilitate discharge, whereas a confirmed fracture can still require treatment, immobilization, consultation, or procedure planning.
An unchanged fracture-confirmed series therefore supports outcome specificity. It does not exclude every concurrent change that affected fracture-excluded visits differently. Nor is the classification completely external to the workflow: fracture status came from treating physicians' ICD-coded discharge diagnoses without study-level imaging re-adjudication. If deployment changes recognition or coding, the composition of the two groups can change with the intervention.
Audit the Outcome, Not Just the Algorithm
Length of stay is operationally meaningful, but it is produced by the whole care system. Arrival volume, staffing, imaging queues, consultation delays, discharge processes, boarding, and case severity can all move it. The evaluation should show how these factors behaved around go-live and whether adjustment decisions were made before inspecting the result.
Distributional results also matter. An eight-minute mean reduction can describe a small improvement for everyone, a major improvement for a delayed minority, or a mixture of benefit and harm. In this example, the authors reported no change near the lower part of the distribution and a larger reduction in its slow tail. That points toward a plausible bottleneck mechanism and tells implementers where to look next.
A Practical Decision Rule for Deployment Claims
- Define the intervention as a workflow. Record the model, interface, alert routing, training, uptake, overrides, and phase-in period.
- Draw the no-deployment trajectory. Use adequate pre-deployment observations and a concurrent comparison when one has a defensible counterfactual role.
- Map every co-intervention. Staffing, protocols, hardware, capacity, and coding changes belong on the same timeline as go-live.
- Stabilize measurement. Apply the same endpoint definition across periods and validate code-based outcomes when the tool can alter diagnosis or documentation.
- Pre-specify the impact shape. Immediate level change, delayed effect, gradual slope change, and temporary disruption are different hypotheses.
- Match the language to the design. “Associated with change after deployment” is not timid when the counterfactual remains partly assumed. It is accurate.
Reviewer Red-Flag Checklist
Why This Matters for Aqrab
AI deployment claims are assembled across a paper: the implementation timeline, comparison logic, outcome definition, analytic impact model, distributional results, and causal language rarely sit in one place. A useful critique reconnects them. It asks whether the paper's counterfactual is as carefully engineered as its algorithm.
Use Aqrab Try to pressure-test the reasoning in a clinical AI deployment paper. The goal is not to reject real-world evidence. It is to identify the strongest claim the design has actually earned.
Methods Sources
- Graif N, et al. AI-assisted radiographic fracture detection and length of stay in the adult ambulatory orthopedic emergency department: a before-after cohort study with a disease-specific internal control. International Journal of Medical Informatics. 2026;221:106679. doi:10.1016/j.ijmedinf.2026.106679.
- Bernal JL, Cummins S, Gasparrini A. Interrupted time series regression for the evaluation of public health interventions: a tutorial. International Journal of Epidemiology. 2017;46(1):348–355.
- Sterne JAC, et al. Assessing risk of bias in a non-randomized study. Cochrane Handbook for Systematic Reviews of Interventions, Chapter 25.
- Harris AD, et al. Research methods in healthcare epidemiology and antimicrobial stewardship: quasi-experimental designs. Infection Control & Hospital Epidemiology. 2016;37(10):1135–1140.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
The Denominator Illusion: Why 20,000 Measurements May Still Mean 200 Patients
A practical unit-of-analysis guide for clinical researchers. Separate rows, patients, clusters, and target populations before repeated observations create false precision.
Ecological Fallacy: When Hospital-Level Data Become Patient-Level Advice
A practical ecological fallacy guide for clinical researchers. Match the unit of analysis, exposure, outcome, and claim before turning group-level associations into patient-level advice.
Crossover Trials: When Every Patient Is Their Own Control—and Their Own Carryover Problem
A practical guide to crossover trials for clinical researchers. Audit treatment reversibility, washout, carryover, period effects, sequence, dropout, and paired analysis before trusting an efficient within-patient comparison.