← Back to Blog
Causal InferenceReal-World EvidenceStudy DesignMethods Critique

Causal Readiness: When a Huge Linked Dataset Still Cannot Identify an Effect

August 4, 2026·13 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A database contains millions of linked health, education, and administrative records. The cohort is enormous, the propensity-score distributions overlap, and the analysis plan names a modern estimator. It is tempting to conclude that the study is ready for causal inference. But scale answers a question about how much data you have. It does not answer whether the data record the clinical decision you want to evaluate.

Causal readiness comes before model selection. Before choosing inverse-probability weighting, augmented estimation, or g-computation, ask whether the dataset can represent the intervention, the comparator, the reasons for assignment, and the outcome well enough to support the target contrast. If one of those pieces is missing, a larger sample may give you a narrower confidence interval around an assumption you cannot defend.

Big Data Can Be Precise About the Wrong Question

Linked administrative data are often assembled for care delivery, payment, education, or service management—not for the causal question a new study asks. A record may show that a service was received without showing its content, intensity, timing, eligibility rule, or the clinical need that led to it. An outcome code may be reliable for one purpose and too coarse for another.

This is not an argument to discard linked data. It is a reason to treat data provenance and measurement as part of identification. A well-specified target trial can expose the gaps. It cannot fill them in by naming a better estimator.

The reviewer’s first question

If this were a randomized trial, what information would the protocol collect that the linked dataset does not?

Four Gates Before the Estimator

Assignment and need

Can you measure the severity, eligibility, preferences, capacity, or other reasons that influenced treatment choice?

Intervention

Is “received treatment” a reproducible strategy, or a label hiding different doses, providers, pathways, or start dates?

Comparator

Were comparison patients eligible for both strategies at the same time zero, or are they simply people who were easy to link?

Outcome

Does the recorded outcome capture the endpoint, timing, and clinical meaning that the estimand requires?

These gates are related but not interchangeable. Good outcome ascertainment cannot repair unmeasured treatment need. Common support cannot repair an intervention definition that mixes incompatible services. A carefully aligned time zero cannot turn a proxy outcome into the outcome the protocol promised.

Causal-readiness gate

Is the data ready for the causal question?

Mark each gate only when the study can defend it with a reproducible variable, time window, or measurement process. This is a design prompt, not a validated score.

Ready for a causal design review

The core data gates are provisionally open. You still need to defend time zero, exchangeability, positivity, missingness, and the target estimand.

Passing these gates does not prove identification. It tells you whether the dataset is ready for the harder assumptions to be examined honestly.

A Linked-Data Example: School Placement and Neurodisability

A recent ECHILD study used linked health and education data to frame a target-trial emulation comparing mainstream with specialist secondary-school placement for pupils with neurodisability. The cohorts were large and the authors reported common support in the propensity-score distributions. That is useful evidence for one necessary design condition: the observed data contain some pupils in both placement categories across the measured covariates.

It is not evidence, by itself, that placement is exchangeable after adjustment. Specialist placement was associated with greater health complexity, absence, and deprivation. Those patterns make the measurement of need and assignment central to the causal question. The study’s planned IPW, augmented IPW, and g-computation analyses are estimators; they are not substitutes for recording the variables that make treatment selection plausible to adjust for.

Do not promote overlap into identification

Common support says both strategies occurred. It does not say the measured covariates capture why they occurred, or that the recorded intervention and outcome are adequate for the target trial.

The Data-Readiness Failure Modes Reviewers Miss

What the paper showsWhat it still has to defend
Millions of linked recordsWhether the missing clinical variables are missing at random, irrelevant, or decisive for assignment.
Propensity-score overlapWhether measured covariates support exchangeability and whether the treatment contrast is clinically coherent.
A treatment indicatorThe content, dose, timing, provider, adherence, switching, and versions represented by that indicator.
A coded outcomeWhether ascertainment is comparable across treatment groups and matches the estimand’s follow-up window.
A sophisticated estimatorWhy its assumptions are more credible than a simpler design, and which assumptions remain untestable.

What to Do When a Gate Is Closed

Improve the measurement

Link to additional need, intervention, or outcome measures when the missing information is plausibly decisive. Document what the linkage adds and what it still cannot observe.

Narrow the question

Replace a vague service category with a strategy the data can actually distinguish. A narrower estimand is more honest than a broad label with hidden heterogeneity.

Use design evidence, not just model evidence

Consider negative controls, falsification outcomes, natural experiments, or an instrumental variable only when the substantive assumptions are defensible—not as decorative extras.

Say what the data cannot answer

A feasibility or descriptive analysis can still be useful. Label it as such instead of letting a precise association inherit a causal interpretation.

Decision Rules for a Busy Review

  • If the treatment is a broad administrative category, ask what clinically different interventions it combines.
  • If the exposure is strongly selected by need, inspect what need is absent before celebrating a large sample or good balance.
  • If common support is reported, treat it as a positivity check—not as proof of exchangeability.
  • If the outcome is a billing, attendance, or utilization code, explain why its measurement is comparable across strategies.
  • If no credible variable or design addresses a closed gate, change the claim before changing the estimator.

Why This Matters for Aqrab

The most expensive methods mistake is often made before the model is fitted: asking a database to answer a question it was never designed to observe. Aqrab helps researchers pressure-check the causal story at that boundary—what was measured, when it was measured, which contrast is actually represented, and which assumptions are still doing the work.

Use Aqrab Try to turn a draft causal question into a structured methods critique. Teams building repeatable review workflows can explore /developers.

Methods Anchors

The data-readiness warning is grounded in the ECHILD case study What is needed to make administrative data research ready for causal inference?. The school-placement example comes from a target trial emulation study using the ECHILD database. These papers motivate the critique; they do not establish that any particular intervention is effective or ineffective.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive