← Back to Blog
Study DesignBiostatisticsMethods Critique

The Denominator Illusion: Why 20,000 Measurements May Still Mean 200 Patients

September 8, 2026·12 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A hospital database contains 20,000 blood-pressure readings from 200 patients. The analysis treats every reading as a separate person, reports a tiny standard error, and describes the study as having a sample size of 20,000. The spreadsheet is large. The independent evidence is not.

Measurements from the same patient share biology, care, devices, and history. Patients treated by the same clinician may share decisions. Clinicians in the same hospital may share protocols. When an analysis counts correlated observations as independent, it commits a unit-of-analysis error—often called pseudoreplication. The usual result is confidence that the design did not earn.

Four Units, Four Different Questions

UnitExampleQuestion it answers
ObservationOne blood-pressure readingWhat is stored in one row?
Experimental or assignment unitPatient—or hospital in a cluster trialWhat was assigned to a strategy?
Analytical unitReading, patient summary, or modeled clusterWhat enters the statistical model?
Unit of inferencePatients eligible for similar careWho does the conclusion describe?

These units can differ legitimately. A longitudinal model may use every reading while estimating an average patient-level treatment effect. A cluster trial may assign hospitals, measure patients, and target patients eligible under similar hospital policies. The error is not having multiple levels. It is analyzing or interpreting them as if the hierarchy did not exist.

More Measurements Are Useful—but They Do Not Clone Patients

Repeated observations can describe trajectories, reduce measurement noise, and reveal when an outcome changes. Two eyes can provide information about eye-specific response. Multiple admissions can describe disease burden. None of this makes observations within a person independent.

The amount of distinct information depends on how strongly observations resemble others in the same cluster, how exposure varies, how many independent clusters exist, and what effect is being estimated. There is no universal conversion from “rows” to an effective sample size. But one rule is reliable: adding correlated rows is not equivalent to adding independently sampled patients or hospitals.

Concrete takeaway

Report the full hierarchy: “20,000 readings from 200 patients treated by 35 clinicians at 8 hospitals.” Never let the row count stand in for the number of independent units.

Three Common Failure Modes

Repeated visits counted as new patients

A standard regression treats 20 visits from one patient like visits from 20 unrelated patients. It ignores within-patient correlation and may also let patients with more visits dominate the estimate. Visit frequency can itself reflect illness severity or surveillance.

Body parts, images, or lesions labeled as the sample size

Eyes, teeth, lesions, patches, and images from one person share causes and measurement conditions. A model that splits images at random can even place near-duplicates from the same patient in training and evaluation sets.

Patients analyzed as individually randomized in a cluster trial

When clinics are randomized, patients within a clinic share allocation and context. Ignoring that clustering can produce confidence intervals that are too narrow and P values that are too small.

A Model Name Is Not a Repair Certificate

Mixed-effects models, generalized estimating equations, cluster-robust standard errors, and participant-level summaries can all be appropriate. The choice depends on the estimand, assignment mechanism, outcome distribution, correlation structure, number and size of clusters, and missingness process.

  • Aggregation can produce one value per patient, but may discard timing and within-patient information.
  • Random-effects models represent selected levels of variation, but an omitted level can still leave dependence unmodeled.
  • Marginal models can target population-average effects, but their uncertainty calculations still need an appropriate number of independent clusters.
  • Cluster-robust standard errors do not automatically work well when the number of clusters is small.

Analysis must also match the scientific comparison. If treatment varies only between eight hospitals, thousands of patient records do not create thousands of independent replications of the hospital-level treatment contrast.

The 60-Second Hierarchy Audit

  • What does one row represent: a patient, visit, eye, lesion, image, clinician, or hospital?
  • Which unit was sampled or assigned to the exposure or intervention?
  • Which observations share a patient, clinician, device, site, family, or time period?
  • At what level does the exposure vary—and at what level is the comparison actually replicated?
  • Which population does the final claim describe?
  • Does the analysis model every dependence that matters, and are there enough independent groups for the planned inference?

Put the answers in the methods section and the abstract. Readers should not have to reverse-engineer whether “n = 20,000” means patients, visits, images, or measurements.

Sources and Evidence Maturity

Evidence note: this guide summarizes established statistical design principles. The correct model and small-sample correction are study-specific and should be prespecified with statistical input; this article does not estimate a treatment effect or recommend clinical care.

Where Aqrab Fits

Aqrab helps reviewers surface mismatches between a paper's row count, sampling structure, assignment level, analysis, and final claim. Try Aqrab with a methods section that contains repeated or clustered data, or use the developer documentation to add a hierarchy check to a structured review workflow.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive