Informative Cluster Size: When the Biggest Sites Start Writing the Result
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
Clustered data are not troublesome only because patients from the same hospital or practice resemble one another. They are also troublesome because some clusters are much bigger than others, and that size difference is often telling you something.
A large referral center may see sicker patients, adopt new treatments faster, measure outcomes more intensely, or keep better records than small sites. If cluster size is tied to prognosis, treatment uptake, or outcome ascertainment, the largest sites do not just improve precision. They can quietly determine which population your treatment effect mostly represents.
The Core Design Rule
If cluster size carries information about the clinical process, ask first whether your target is the average patient, the average center, or a specific deployment mix of centers. The weighting choice is part of the estimand, not a technical afterthought.
Decision rule:
Do not let "we accounted for clustering" end the conversation. First ask whether the analysis answered the patient-average question, the center-average question, or an unlabeled hybrid driven by whoever enrolled the most people.
Why Informative Cluster Size Is a Design Problem
Big sites are often different sites
Volume can proxy referral severity, specialist availability, algorithmic triage, or richer data capture. Size often travels with structure.
Standard cluster fixes solve a narrower problem
Robust standard errors, mixed models, and GEEs handle correlation. They do not automatically tell you whether the weighting behind the effect estimate matches the question you meant to ask.
The drift can change the story
If high-volume centers have smaller benefits, larger harms, or simply different treatment uptake, the pooled answer may mostly describe them while sounding universal.
A Concrete Clinical Example
Case
A multicenter sepsis pathway study where tertiary hospitals dominate enrollment
Imagine a multicenter comparative-effectiveness study of an early sepsis treatment pathway. A handful of tertiary centers contribute most patients, treat more aggressively, and receive the sickest referrals. Smaller hospitals enroll fewer patients and start treatment later, but their baseline risk is lower.
A patient-level model with site adjustment may be entirely legitimate if the question is "what is the average effect across enrolled patients?" But if the manuscript speaks as if it estimated the average effect across hospitals, that interpretation has already drifted. The biggest centers wrote most of the answer.
This matters in cluster randomized trials too. Unequal practice size can mean a treatment policy is judged mostly by large practices, even if the implementation decision will later be spread across many smaller sites.
Interactive cluster-size explorer
When the biggest sites carry most of the patients, they can quietly carry most of the answer too
This toy example compares one small site type and one large site type. Change the site sizes, treatment mix, and outcome risks to see how a patient-weighted answer can drift away from the simple average of site-level effects.
Larger centers dominate the patient-level answer even before the regression model begins speaking.
Informative cluster size often arrives together with different treatment adoption patterns across sites.
Bigger sites can be referral centers with sicker patients, or the reverse. Either way, site size may map to baseline prognosis.
Large-site share
84.6%
This is the share of all patients contributed by the large-site stratum.
Small-site effect
-5.0 percentage points
Large-site effect
-3.0 percentage points
Patient-weighted effect
-1.2 percentage points
| Quantity | Approximate value | Why it matters |
|---|---|---|
| Center-average effect | -4.0 percentage points | This is the simple average of the small-site and large-site treatment effects. |
| Patient-weighted effect | -1.2 percentage points | This is the answer produced when bigger sites contribute more patients and therefore more leverage. |
| Estimand drift | 2.8 percentage points | If this moves meaningfully, ask whether the manuscript described the target population honestly. |
This is a teaching illustration, not a validated estimator. Its job is to make the weighting question visible before "cluster-robust" starts sounding like a full design solution.
What the Weighting Is Really Choosing
| If you want to learn about... | Then the analysis should... | What goes wrong when it does not |
|---|---|---|
| Average effect across all patients who could enter the study | A patient-level estimand can be reasonable, but state openly that larger sites will contribute more to the answer. | If site size tracks prognosis or treatment adoption, the result may mostly describe the environments that enroll the most patients. |
| Average effect across centers as decision units | Use center-level summaries or methods whose weighting reflects a center-average estimand. | Treating this as a patient-average result can flatten meaningful between-center heterogeneity. |
| Policy question about a future health system with a known site mix | Prespecify that mix and justify why those weights match the deployment setting. | Post hoc weighting can become a story-telling device rather than a design choice. |
What Reviewers Should Ask Before Trusting the Result
- Did the authors say explicitly whether the estimand is patient-average, center-average, or policy-weighted?
- Does cluster size plausibly reflect disease severity, treatment access, referral patterns, or outcome capture?
- Would the headline effect change if centers contributed equally rather than in proportion to their patient counts?
- Are treatment uptake and baseline risk different across large versus small sites in ways that make size informative?
- Did the authors separate the correlation problem from the weighting problem, or did "cluster-adjusted" try to cover both?
Reviewer Red Flags
- The paper uses cluster-robust standard errors or a random intercept, but never explains what population the effect represents.
- Large sites contribute most treated patients and have visibly different prognosis, workflow, or case mix from smaller sites.
- Site size is linked to referral intensity, specialist staffing, data completeness, or treatment adoption, but the discussion treats cluster size as harmless bookkeeping.
- A center-level summary and a patient-level summary would answer different questions, yet only one is shown and no sensitivity analysis explores the gap.
- The manuscript says it adjusted for site without clarifying whether site was a nuisance correlation structure or part of the causal design problem.
The Practical Bottom Line
Informative cluster size is the kind of problem that hides in plain sight. The methods section sounds mature because the authors mention clustering, random effects, or robust variance. But the harder question is quieter: whose experience does the pooled estimate mostly represent?
If you are reviewing a multicenter registry, cluster randomized trial, or AI-generated methods summary that glosses over this distinction, Aqrab can help stress-test the estimand, weighting choice, and reviewer red flags before a polished site-adjusted model starts sounding more universal than it is. If you want to build those checks into your own review workflow, the developer tools are the right place to begin.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Run-In Periods: When Your Trial Randomizes the Easy Patients First
A practical guide to run-in periods for clinical researchers. Covers adherence enrichment, tolerability selection, estimand drift, external validity, and what reviewers should demand before trusting a polished randomized cohort.
Bayesian Borrowing: When Historical Data Starts Spending Credibility It Did Not Earn
A practical guide to Bayesian borrowing for clinical researchers. Covers exchangeability, commensurate priors, historical controls, calendar-time drift, and what reviewers should demand before trusting extra certainty borrowed from earlier data.
Baseline Adjustment in Randomized Trials: Why Change From Baseline Keeps Losing to ANCOVA
A practical guide to baseline adjustment in randomized trials. Covers ANCOVA versus change scores, percent change traps, responder thresholds, and what reviewers should demand before trusting a tidy efficacy claim.