Specification Curve Analysis: When One Model Hides a Vibration of Effects
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Specification curve analysis asks a question that one preferred model cannot answer: how much does the conclusion move across the analytic choices a competent researcher could reasonably have made? In observational clinical research, decisions about outcome definitions, eligibility, missing data, adjustment, and follow-up can turn one dataset into many defensible analyses—and sometimes opposing claims.
The goal is not to replace judgment with every model a computer can fit. It is to make judgment visible. A useful specification curve defines the reasonable analysis set before inspecting results, estimates every member consistently, and shows which decisions make the scientific conclusion stable or fragile.
The Clean Metaphor: Show the Contact Sheet
A preferred analysis is one photograph. A specification curve is the contact sheet.
The selected frame may be sharp and honest. The contact sheet reveals whether nearby, equally defensible frames tell the same story—or whether the apparent finding depends on one convenient crop.
Interactive specification explorer
One Estimate or the Whole Contact Sheet?
This simulated study compares two treatments on an absolute event-risk scale. Reveal the full set of prespecified, defensible analyses. The numbers teach interpretation; they are not results from a real study.
The preferred model looks decisive
It estimates 18 fewer events per 1,000 and its interval excludes no difference. That may be a defensible primary result, but it does not reveal whether equally defensible choices tell the same story.
Reviewer lesson: do not elect a winner by majority vote. Ask which choices change the estimand, which probe assumptions about the same estimand, and why the preferred analysis deserves priority.
What Belongs in a Defensible Specification Universe?
Imagine an electronic-health-record study comparing two antihypertensive drugs on acute kidney injury. Reasonable uncertainty may remain about whether to use a narrow validated outcome code or a broader code-plus-laboratory definition, how to handle missing baseline kidney function, and whether to restrict to patients with clinical overlap between treatment groups. Those choices can be prespecified and displayed rather than hidden inside a supplementary table.
| Choice | Include when | Do not pretend |
|---|---|---|
| Outcome algorithm | More than one clinically credible definition exists | An unvalidated broad code is equivalent to adjudication |
| Missing-data strategy | Assumptions differ but remain scientifically plausible | Complete cases and imputation target the same population automatically |
| Adjustment set | Each set follows a defensible causal structure | Every measured variable is a candidate confounder |
| Follow-up window | Each window answers a clinically relevant timing question | A 30-day and 1-year effect are one estimand |
Robustness Checks and Estimand Changes Are Not the Same
Some specifications probe assumptions about one target effect. Alternative valid outcome algorithms, models for missing covariates, or estimators of the same contrast may fit this role. Other choices alter the target itself. Changing from all eligible patients to overlap-restricted patients changes the population. Changing follow-up from 30 days to one year changes the time horizon. Switching from an intention-to-treat effect to a per-protocol effect changes the treatment strategy.
Reviewer move
Group specifications by estimand before comparing them. Instability within one estimand challenges robustness. Differences across estimands may be clinically expected and should be explained, not averaged away.
How to Read a Specification Curve Without Taking a Vote
- Start with the target question. Confirm that the population, treatment strategies, outcome, time horizon, and effect measure are explicit.
- Inspect the universe. Ask whether all specifications were scientifically defensible and defined without looking at favorable results.
- Read magnitude and uncertainty. A pile of p-values hides whether estimates are clinically similar, imprecise, or opposite in direction.
- Find the hinges. Identify which choice—outcome, population, adjustment, missingness, or follow-up—moves the estimate most.
- Locate the preferred model. If it sits at an extreme, the protocol needs a strong scientific reason for privileging it.
- Keep causal validity separate. Agreement across many models cannot repair immortal time, unmeasured confounding, selection bias, or a wrong comparator shared by all of them.
The specifications are also correlated because they reuse the same records. Twelve similar models are not twelve independent replications. Formal specification-curve inference exists, but the clinical audit still begins with whether the analysis universe and its assumptions make scientific sense.
Three Ways the Method Can Become Automated P-Hacking
First, authors can define “reasonable” after seeing the curve, excluding inconvenient specifications until the story stabilizes. Second, they can flood the universe with trivial variants of one favorable model, allowing a model count to masquerade as evidence. Third, they can include causally indefensible adjustment sets and then celebrate the median estimate. Exhaustiveness is not validity.
The safeguard is a protocol-level specification map: list each decision, its allowed options, its causal or measurement rationale, the estimand it targets, and the rule for privileging any primary analysis. Code should generate the grid from that map, not invent the map from the output.
Reviewer Red-Flag Checklist
Why This Matters for Aqrab
Analytic discretion is scattered across a paper: cohort construction in one section, outcome coding in another, adjustment in a table, missingness in a footnote, and sensitivity analyses in a supplement. A rigorous critique reconnects those choices and asks whether the headline survives the full analysis path rather than merely one final model.
Use Aqrab Try to pressure-test the specification logic in an observational study. The practical question is not “Did the authors run many models?” It is “Did they expose the decisions that could have changed the claim?”
Methods Sources
- Patel CJ, Burford B, Ioannidis JPA. Assessment of vibration of effects due to model specification can demonstrate the instability of observational associations. Journal of Clinical Epidemiology. 2015;68(9):1046–1058. doi:10.1016/j.jclinepi.2015.05.029.
- Steegen S, Tuerlinckx F, Gelman A, Vanpaemel W. Increasing transparency through a multiverse analysis. Perspectives on Psychological Science. 2016;11(5):702–712. doi:10.1177/1745691616658637.
- Simonsohn U, Simmons JP, Nelson LD. Specification curve analysis. Nature Human Behaviour. 2020;4:1208–1214. doi:10.1038/s41562-020-0912-z.
- Tierney BT, et al. Leveraging vibration of effects analysis for robust discovery in observational biomedical data science. PLOS Biology. 2021;19(9):e3001398. doi:10.1371/journal.pbio.3001398.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Delayed Entry in Survival Analysis: Why Nobody Is at Risk Before They Enter
A practical guide to left truncation in clinical survival analysis. Align the origin, entry time, event time, and risk set before interpreting Kaplan–Meier curves or Cox models from prevalent cohorts.
Target Trial Emulation Cannot Randomize Clinical Judgment: A Pertussis Study Audit
A practical target-trial emulation audit using an infant pertussis study. Check clinical-judgment confounding, propensity-score overlap, endpoint timing, sparse outcomes, and claim strength.
Vaccine Effectiveness Without Matching: Why Calendar Time Comes Before Pairing
A practical guide to calendar time in vaccine-effectiveness studies. Learn how changing uptake and infection hazards alter risk sets, estimands, and target-trial conclusions.