← Back to Blog
Missing DataClinical TrialsMethods Critique

Complete-Case Analysis: When Missing Data Quietly Changes the Study Population

July 7, 2026·15 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Complete-case analysis survives because it feels disciplined. The analyst refuses to guess, keeps only rows with observed values, and proceeds with what looks like a clean dataset. The problem is that clean is not the same thing as representative.

Once missingness is related to prognosis, treatment tolerability, follow-up intensity, or survival, the complete cases stop being a neutral subset of the study. They become the patients who were easiest to observe. At that moment, the method is no longer just losing power. It is changing the population and often changing the question.

The Core Decision Rule

Do not ask whether complete-case analysis is simple. Ask whether the people with observed endpoints are still exchangeable with the people who disappeared.

Decision rule:

If becoming a complete case depends on post-baseline health, treatment response, adverse events, site behavior, or any predictor of the outcome, treat complete-case analysis as a selection mechanism first and an estimation method second.

The rare low-risk setting is administrative missingness that is both uncommon and plausibly unrelated to the unobserved outcome after conditioning on variables already in the analysis. Most manuscripts do not earn that claim by saying the missing proportion was modest.

Why Researchers Reach for It Anyway

It feels assumption-light

Analysts hear “no imputation” and imagine they avoided assumptions. In reality, they assumed the observed subset answers the same question as the full cohort.

Software makes it effortless

Many pipelines silently drop incomplete rows before the analyst has even decided whether the resulting population still matches the estimand.

Tables still look respectable

A neat baseline table for completers can hide that the sickest, least adherent, or least reachable patients are no longer being asked to vote on the treatment effect.

A Concrete Clinical Example

Case

A pragmatic hypertension trial with more missed endpoint visits after treatment side effects

Imagine a 6-month pragmatic trial of a new antihypertensive strategy. The treatment arm lowers blood pressure well among patients who stay engaged, but dizziness and medication changes lead some frailer patients to miss the final clinic visit. In usual care, fewer patients miss the endpoint.

A complete-case analysis compares the patients with observed 6-month blood pressure and reports a strong benefit. But the missing treated patients were not random administrative accidents. They were the patients whose symptoms, treatment changes, or fragmented follow-up made the outcome harder to capture and probably less favorable.

The estimate may still be numerically precise. It is just no longer the effect in the randomized population. It is the effect in the subset who tolerated the pathway and remained measurable.

Interactive complete-case explorer

When the patients who vanish are doing worse, the clean rows can flatter the treatment

This teaching toy uses a continuous outcome where higher values are better at the endpoint. The complete-case estimate compares only observed patients. The full-data estimate adds back the patients whose endpoint would have been worse or better than the observed group.

Selection distortion5.0 pointscomplete-case effect minus full-data effect

Differential missingness gives complete-case analysis more room to replace the randomized cohort with a selected survivor subset.

Balanced missingness is not automatically safe, but imbalance is often the first visible warning sign.

Lower values represent patients who left because of toxicity, relapse, nonresponse, or worsening frailty.

If missing control patients look different from missing treated patients, one shared missing-data story is already too simple.

Rows retained

80.0%

This is the share of the randomized or enrolled cohort that survives into the complete-case analysis.

Complete-case effect

8.0 points

This is what the manuscript would report if it compared only patients with observed endpoints.

Full-data effect

3.0 points

This is the effect after the missing patients are allowed to have their own outcome trajectory.

QuantityApproximate valueWhy it matters
Treatment mean after restoring missing patients62.4 pointsIf missing treated patients were doing worse than completers, the apparent treatment performance falls.
Control mean after restoring missing patients59.4 pointsComplete-case bias depends on both arms. A messy control arm can distort the answer just as easily.
Selection distortion5.0 pointsPositive values mean complete-case analysis is making the treatment look better than the full cohort would justify.

This is a teaching illustration, not a validated estimator. It exists to show why complete-case analysis is a population-selection decision before it is a software checkbox.

When Complete-Case Analysis Most Commonly Fails

Failure modeWhat gets selectedWhy the estimate drifts
Adverse-event dropoutPatients who tolerated treatment better stay visible.The analysis can exaggerate efficacy by keeping the beneficiaries and losing the harmed.
Informative visit attendancePatients with more clinic contact contribute more final outcomes.Completeness starts tracking engagement, site behavior, and disease severity rather than random measurement.
Real-world covariate missingnessPatients with richer documentation stay in the model.The “analyzable” cohort can become younger, more connected to care, and less clinically chaotic.
Post-outcome conditioningOnly survivors or patients stable enough to be measured remain.The method quietly conditions on downstream health status, which can create collider structures and estimand drift.

When It May Be Defensible

Complete-case analysis is not automatically forbidden. It is just rarely innocent. A defensible use case usually needs all of the following:

1. Missingness is mostly administrative

Examples include a laboratory platform outage, a delayed shipment, or a site closure that affected outcome capture without obvious linkage to the patient's likely endpoint value.

2. The missing share is small and balanced

This does not prove safety, but large differential missingness should end the conversation before it starts pretending to be a minor detail.

3. Sensitivity analyses tell the same story

If multiple imputation, inverse-probability weighting for observation, or plausible MNAR scenarios materially change the result, the complete-case answer was never stable enough to headline.

Reviewer Red Flags to Demand Before Trusting the Estimate

1. The manuscript reports only the percentage missing

Proportion missing is not a mechanism. Readers need reasons for missingness, timing, and whether missingness differed by arm, site, severity, or outcome trajectory.

2. Completers are described, but noncompleters are invisible

If you never see who disappeared, you cannot judge how badly the analysis population drifted.

3. Post-randomization reasons for missingness are treated like baseline noise

Toxicity, rescue treatment, hospitalization, and symptom worsening are not benign excuses for incomplete data. They are part of the clinical story.

4. No sensitivity analysis competes with the complete-case result

A single tidy estimate should not get the last word when the missing-data assumptions were never stress-tested.

What a Better Analysis Looks Like

The right fix depends on why data are missing and what outcome is being estimated. Sometimes multiple imputation is reasonable. Sometimes weighting the observation process is the better diagnostic frame. Sometimes the honest move is to show that plausible MNAR scenarios break the claim. The point is not that one replacement method always wins. The point is that complete-case analysis should have to defend itself against alternatives, not be treated as the default because it is easy.

If your team is trying to pressure-test missing-data decisions before they harden into analysis code, Aqrab is built for exactly that kind of methodological review. The fastest path is to start in Aqrab Try and make the missing-data assumptions explicit while the study still has room to change.

The Bottom Line

Complete-case analysis is not just a smaller dataset. It is often a filtered dataset whose filters were applied by clinical reality after the study began. When missingness follows prognosis, treatment tolerance, or follow-up intensity, the method stops estimating the answer you thought you randomized or designed the study to learn.

The cleanest sentence to keep in mind is this: missing rows are part of the result, not housekeeping around the result.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive