← Back to Blog
Study DesignCausal InferenceMethods Critique

Before the Model: A Five-Gate Study Design Audit for Clinical Research

August 6, 2026·12 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A research team spends weeks debating the regression family, tunes a machine-learning pipeline, and produces a beautifully formatted estimate. Then a reviewer asks a short question: what exactly was this study designed to identify? The team discovers that treatment started at different times, the comparison group came from a different care pathway, and the primary outcome was not measured the same way for everyone.

The problem is not that the analysis was too simple. It is that the design never made the intended claim identifiable. Statistical sophistication can reduce variance, model nonlinearities, and handle complex data. It cannot turn a vague comparison into a well-defined intervention, repair a post-treatment variable, or make a selected sample represent a population it never had a chance to enter.

Design Is the Part That Decides What the Number Means

“Study design” is not just the label on the first page. It is the set of decisions that connect a question to an observable comparison: who is eligible, when follow-up begins, what counts as exposure, what outcome is measured, and which departures from the plan change the question.

A randomized trial gets important help from allocation, but it still needs an estimand, outcome definition, follow-up window, and plan for rescue therapy or treatment switching. An observational study has the same design obligations with less protection from confounding. A prediction study has a different target, but it still must define the prediction time and prevent the model from seeing information that would arrive later.

The design-first rule

Before asking “Which model should we use?”, ask “Which population, decision, contrast, outcome, and time horizon would make this estimate answerable?”

Interactive five-gate design audit

Find the bottleneck before you choose the model

Pick the question, data setting, and timing that best describe your project. This teaching tool surfaces the first design decision to pressure-test. It is a triage aid, not a substitute for a protocol or statistical analysis plan.

First readModel-ready only after design checks

Your setup

  • Question: Causal effect: what would change under an intervention?
  • Setting: Observational clinical data
  • Timing: Features defined before the decision point

An observational estimate needs an explicit treatment strategy, comparator, time zero, confounding set, and outcome window before adjustment can mean anything.

Next audit: Draw the causal structure and test overlap, treatment measurement, outcome capture, and follow-up before fitting the effect model.

Ask the team: What assumptions would make the observed comparison exchangeable enough for the causal contrast you want?

The combinations are intentionally simplified. Real readiness depends on the target estimand, data-generating process, measurement quality, missingness, and the assumptions your study can defend.

The Five Gates

Use these gates in order. A failure at an earlier gate can make later statistical refinement beside the point.

1. Question and estimand

Is the aim causal, predictive, descriptive, or something else? For a causal claim, state the intervention or treatment strategy, target population, outcome, time horizon, and contrast. “The association between X and Y” is not yet a causal estimand.

2. Population and time zero

Who could enter, and when does eligibility become true? Time zero should be a real decision point shared by the comparison, not the first convenient timestamp in the database. If the treated group must survive or remain event-free until classification, the design may be granting it immortal time.

3. Comparison and treatment version

What exactly is being compared? “Received the drug” may mix dose, timing, formulation, titration, monitoring, adherence support, and background care. A control group can also be clinically different before analysis begins. If the strategies cannot be written into a protocol, the contrast is probably still underspecified.

4. Outcome and measurement

Is the outcome patient-important, measured in the same way, and captured within the promised window? A code, score, biomarker, or composite can be useful, but the meaning may vary by site, treatment, rater, device, or surveillance intensity. Precision does not make a noisy or differentially measured outcome valid.

5. Follow-up, missingness, and deviations

What happens after the decision? Rescue therapy, switching, loss to follow-up, competing events, and missing outcomes are not interchangeable nuisances. Each one may imply a different estimand. Decide whether it is part of the treatment policy, a censoring event, an outcome, or a reason to change the question.

What Sophisticated Analysis Can and Cannot Do

ProblemWhat analysis may help withWhat it cannot assume away
Measured confoundingAdjustment, weighting, standardization, or a doubly robust estimatorUnmeasured need, poor treatment measurement, or no overlap
Repeated or clustered observationsMixed models, marginal models, or cluster-robust inferenceA target population or exposure that was never defined
Missing outcomesImputation and sensitivity analyses under stated assumptionsA missingness process that was not measured or plausibly modeled
Prediction performanceCalibration, discrimination, validation, and decision-curve analysisA treatment effect or a deployment setting different from validation

How to Make a Design Review Concrete

Before analysis begins, ask the team to produce five short artifacts:

  • One-sentence estimand: “Among [population], what is the effect of [strategy A] versus [strategy B] on [outcome] by [time]?”
  • Eligibility and time-zero table: show which information is known before assignment and when follow-up begins.
  • Exposure and comparator protocol: include dose, timing, versions, cointerventions, and allowed deviations.
  • Outcome measurement map: define the source, ascertainment window, adjudication, and likely differential detection.
  • Threat-to-claim register: for each assumption, record the evidence, diagnostic, sensitivity analysis, and what conclusion would be softened if it failed.

This is not bureaucracy for its own sake. It makes the design inspectable before the model creates a false sense of completion. A reviewer can disagree with an assumption that is visible. A hidden assumption can only become a late-stage surprise.

Why This Matters for Aqrab

Aqrab is most useful before a study is polished enough to hide its weak points. Its job is not to reward a method name. It is to pressure-test whether the question, data, comparison, measurement, and assumptions line up well enough for the proposed claim.

Paste a protocol, methods section, or research question into Aqrab Try for a structured critique. If you are building a repeatable review workflow, explore the developer tools to put the same design-first questions upstream of analysis.

Methods Anchors

The causal-design framing follows the target-trial approach described by Hernán and Robins. The design-first review sequence also aligns with the reporting questions emphasized by the STROBE guidance for observational studies and the CONSORT guidance for randomized trials. These frameworks improve transparency; they do not prove that every identification assumption is true.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive