The Denominator Illusion: Why 20,000 Measurements May Still Mean 200 Patients
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
A hospital database contains 20,000 blood-pressure readings from 200 patients. The analysis treats every reading as a separate person, reports a tiny standard error, and describes the study as having a sample size of 20,000. The spreadsheet is large. The independent evidence is not.
Measurements from the same patient share biology, care, devices, and history. Patients treated by the same clinician may share decisions. Clinicians in the same hospital may share protocols. When an analysis counts correlated observations as independent, it commits a unit-of-analysis error—often called pseudoreplication. The usual result is confidence that the design did not earn.
Four Units, Four Different Questions
| Unit | Example | Question it answers |
|---|---|---|
| Observation | One blood-pressure reading | What is stored in one row? |
| Experimental or assignment unit | Patient—or hospital in a cluster trial | What was assigned to a strategy? |
| Analytical unit | Reading, patient summary, or modeled cluster | What enters the statistical model? |
| Unit of inference | Patients eligible for similar care | Who does the conclusion describe? |
These units can differ legitimately. A longitudinal model may use every reading while estimating an average patient-level treatment effect. A cluster trial may assign hospitals, measure patients, and target patients eligible under similar hospital policies. The error is not having multiple levels. It is analyzing or interpreting them as if the hierarchy did not exist.
More Measurements Are Useful—but They Do Not Clone Patients
Repeated observations can describe trajectories, reduce measurement noise, and reveal when an outcome changes. Two eyes can provide information about eye-specific response. Multiple admissions can describe disease burden. None of this makes observations within a person independent.
The amount of distinct information depends on how strongly observations resemble others in the same cluster, how exposure varies, how many independent clusters exist, and what effect is being estimated. There is no universal conversion from “rows” to an effective sample size. But one rule is reliable: adding correlated rows is not equivalent to adding independently sampled patients or hospitals.
Concrete takeaway
Report the full hierarchy: “20,000 readings from 200 patients treated by 35 clinicians at 8 hospitals.” Never let the row count stand in for the number of independent units.
Three Common Failure Modes
Repeated visits counted as new patients
A standard regression treats 20 visits from one patient like visits from 20 unrelated patients. It ignores within-patient correlation and may also let patients with more visits dominate the estimate. Visit frequency can itself reflect illness severity or surveillance.
Body parts, images, or lesions labeled as the sample size
Eyes, teeth, lesions, patches, and images from one person share causes and measurement conditions. A model that splits images at random can even place near-duplicates from the same patient in training and evaluation sets.
Patients analyzed as individually randomized in a cluster trial
When clinics are randomized, patients within a clinic share allocation and context. Ignoring that clustering can produce confidence intervals that are too narrow and P values that are too small.
A Model Name Is Not a Repair Certificate
Mixed-effects models, generalized estimating equations, cluster-robust standard errors, and participant-level summaries can all be appropriate. The choice depends on the estimand, assignment mechanism, outcome distribution, correlation structure, number and size of clusters, and missingness process.
- Aggregation can produce one value per patient, but may discard timing and within-patient information.
- Random-effects models represent selected levels of variation, but an omitted level can still leave dependence unmodeled.
- Marginal models can target population-average effects, but their uncertainty calculations still need an appropriate number of independent clusters.
- Cluster-robust standard errors do not automatically work well when the number of clusters is small.
Analysis must also match the scientific comparison. If treatment varies only between eight hospitals, thousands of patient records do not create thousands of independent replications of the hospital-level treatment contrast.
The 60-Second Hierarchy Audit
- What does one row represent: a patient, visit, eye, lesion, image, clinician, or hospital?
- Which unit was sampled or assigned to the exposure or intervention?
- Which observations share a patient, clinician, device, site, family, or time period?
- At what level does the exposure vary—and at what level is the comparison actually replicated?
- Which population does the final claim describe?
- Does the analysis model every dependence that matters, and are there enough independent groups for the planned inference?
Put the answers in the methods section and the abstract. Readers should not have to reverse-engineer whether “n = 20,000” means patients, visits, images, or measurements.
Sources and Evidence Maturity
- Altman and Bland, “Units of analysis” — a clinical statistics note explaining how multiple observations per patient can inflate the apparent sample size and why analysis must account for multiplicity.
- Cochrane Handbook, Chapter 23 — official guidance on clustered trials, repeated observations, body parts, crossover designs, and unit-of-analysis errors.
- CONSORT extension for cluster randomized trials — reporting guidance for trials in which groups rather than individuals are randomized.
Evidence note: this guide summarizes established statistical design principles. The correct model and small-sample correction are study-specific and should be prespecified with statistical input; this article does not estimate a treatment effect or recommend clinical care.
Where Aqrab Fits
Aqrab helps reviewers surface mismatches between a paper's row count, sampling structure, assignment level, analysis, and final claim. Try Aqrab with a methods section that contains repeated or clustered data, or use the developer documentation to add a hierarchy check to a structured review workflow.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Ecological Fallacy: When Hospital-Level Data Become Patient-Level Advice
A practical ecological fallacy guide for clinical researchers. Match the unit of analysis, exposure, outcome, and claim before turning group-level associations into patient-level advice.
Crossover Trials: When Every Patient Is Their Own Control—and Their Own Carryover Problem
A practical guide to crossover trials for clinical researchers. Audit treatment reversibility, washout, carryover, period effects, sequence, dropout, and paired analysis before trusting an efficient within-patient comparison.
Causal Identification: Why Longitudinal Data Still Need a Design
A practical causal identification guide for longitudinal clinical research. Learn why measurement order does not create comparability, what repeated observations add, and how to audit causal claims with five reviewer questions.