Before the Model: A Five-Gate Study Design Audit for Clinical Research
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
A research team spends weeks debating the regression family, tunes a machine-learning pipeline, and produces a beautifully formatted estimate. Then a reviewer asks a short question: what exactly was this study designed to identify? The team discovers that treatment started at different times, the comparison group came from a different care pathway, and the primary outcome was not measured the same way for everyone.
The problem is not that the analysis was too simple. It is that the design never made the intended claim identifiable. Statistical sophistication can reduce variance, model nonlinearities, and handle complex data. It cannot turn a vague comparison into a well-defined intervention, repair a post-treatment variable, or make a selected sample represent a population it never had a chance to enter.
Design Is the Part That Decides What the Number Means
“Study design” is not just the label on the first page. It is the set of decisions that connect a question to an observable comparison: who is eligible, when follow-up begins, what counts as exposure, what outcome is measured, and which departures from the plan change the question.
A randomized trial gets important help from allocation, but it still needs an estimand, outcome definition, follow-up window, and plan for rescue therapy or treatment switching. An observational study has the same design obligations with less protection from confounding. A prediction study has a different target, but it still must define the prediction time and prevent the model from seeing information that would arrive later.
The design-first rule
Before asking “Which model should we use?”, ask “Which population, decision, contrast, outcome, and time horizon would make this estimate answerable?”
Interactive five-gate design audit
Find the bottleneck before you choose the model
Pick the question, data setting, and timing that best describe your project. This teaching tool surfaces the first design decision to pressure-test. It is a triage aid, not a substitute for a protocol or statistical analysis plan.
Your setup
- Question: Causal effect: what would change under an intervention?
- Setting: Observational clinical data
- Timing: Features defined before the decision point
An observational estimate needs an explicit treatment strategy, comparator, time zero, confounding set, and outcome window before adjustment can mean anything.
Next audit: Draw the causal structure and test overlap, treatment measurement, outcome capture, and follow-up before fitting the effect model.
Ask the team: What assumptions would make the observed comparison exchangeable enough for the causal contrast you want?
The combinations are intentionally simplified. Real readiness depends on the target estimand, data-generating process, measurement quality, missingness, and the assumptions your study can defend.
The Five Gates
Use these gates in order. A failure at an earlier gate can make later statistical refinement beside the point.
1. Question and estimand
Is the aim causal, predictive, descriptive, or something else? For a causal claim, state the intervention or treatment strategy, target population, outcome, time horizon, and contrast. “The association between X and Y” is not yet a causal estimand.
2. Population and time zero
Who could enter, and when does eligibility become true? Time zero should be a real decision point shared by the comparison, not the first convenient timestamp in the database. If the treated group must survive or remain event-free until classification, the design may be granting it immortal time.
3. Comparison and treatment version
What exactly is being compared? “Received the drug” may mix dose, timing, formulation, titration, monitoring, adherence support, and background care. A control group can also be clinically different before analysis begins. If the strategies cannot be written into a protocol, the contrast is probably still underspecified.
4. Outcome and measurement
Is the outcome patient-important, measured in the same way, and captured within the promised window? A code, score, biomarker, or composite can be useful, but the meaning may vary by site, treatment, rater, device, or surveillance intensity. Precision does not make a noisy or differentially measured outcome valid.
5. Follow-up, missingness, and deviations
What happens after the decision? Rescue therapy, switching, loss to follow-up, competing events, and missing outcomes are not interchangeable nuisances. Each one may imply a different estimand. Decide whether it is part of the treatment policy, a censoring event, an outcome, or a reason to change the question.
What Sophisticated Analysis Can and Cannot Do
| Problem | What analysis may help with | What it cannot assume away |
|---|---|---|
| Measured confounding | Adjustment, weighting, standardization, or a doubly robust estimator | Unmeasured need, poor treatment measurement, or no overlap |
| Repeated or clustered observations | Mixed models, marginal models, or cluster-robust inference | A target population or exposure that was never defined |
| Missing outcomes | Imputation and sensitivity analyses under stated assumptions | A missingness process that was not measured or plausibly modeled |
| Prediction performance | Calibration, discrimination, validation, and decision-curve analysis | A treatment effect or a deployment setting different from validation |
How to Make a Design Review Concrete
Before analysis begins, ask the team to produce five short artifacts:
- One-sentence estimand: “Among [population], what is the effect of [strategy A] versus [strategy B] on [outcome] by [time]?”
- Eligibility and time-zero table: show which information is known before assignment and when follow-up begins.
- Exposure and comparator protocol: include dose, timing, versions, cointerventions, and allowed deviations.
- Outcome measurement map: define the source, ascertainment window, adjudication, and likely differential detection.
- Threat-to-claim register: for each assumption, record the evidence, diagnostic, sensitivity analysis, and what conclusion would be softened if it failed.
This is not bureaucracy for its own sake. It makes the design inspectable before the model creates a false sense of completion. A reviewer can disagree with an assumption that is visible. A hidden assumption can only become a late-stage surprise.
Why This Matters for Aqrab
Aqrab is most useful before a study is polished enough to hide its weak points. Its job is not to reward a method name. It is to pressure-test whether the question, data, comparison, measurement, and assumptions line up well enough for the proposed claim.
Paste a protocol, methods section, or research question into Aqrab Try for a structured critique. If you are building a repeatable review workflow, explore the developer tools to put the same design-first questions upstream of analysis.
Methods Anchors
The causal-design framing follows the target-trial approach described by Hernán and Robins. The design-first review sequence also aligns with the reporting questions emphasized by the STROBE guidance for observational studies and the CONSORT guidance for randomized trials. These frameworks improve transparency; they do not prove that every identification assumption is true.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Causal Identification: Why Longitudinal Data Still Need a Design
A practical causal identification guide for longitudinal clinical research. Learn why measurement order does not create comparability, what repeated observations add, and how to audit causal claims with five reviewer questions.
Causal Readiness: When a Huge Linked Dataset Still Cannot Identify an Effect
A practical guide to causal readiness in linked health and administrative data. Learn why scale and propensity-score overlap are not enough when treatment, need, comparators, or outcomes are poorly measured.
Triangulation in Clinical Research: When One Elegant Design Still Leaves the Same Blind Spot
A practical guide to triangulation in clinical research. Covers what counts as genuinely complementary evidence, how to map designs to specific threats, and what reviewers should demand before trusting “robustness” claims.