← Back to Blog
Clinical TrialsStudy DesignMethods Critique

Small Number of Clusters: When 500 Patients Still Behave Like 10 Sites

July 30, 2026·14 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A cluster-randomized trial can enroll 500 patients and still have the inferential personality of a small study. If those patients come from 10 clinics, the treatment was assigned 10 times, not 500 times. The people inside a clinic share clinicians, protocols, referral patterns, and the same treatment environment. They are not 500 fresh tosses of the randomization coin.

This is the small-number-of-clusters problem. It is easy to miss because the patient count looks impressive and the analysis table may contain a familiar mixed model or a cluster-robust standard error. But clustering creates two separate questions: how much information is lost because observations within a site resemble one another, and whether the inferential method behaves well when the number of sites is small. A method can address the first and still fail the second.

Patients are not the same thing as independent evidence

For equal-sized clusters, a common teaching approximation is the design effect:

design effect = 1 + (average cluster size − 1) × ICC

ICC is the intracluster correlation: how similar two patients from the same cluster are, on average.

If 10 clinics each contribute 50 patients and the ICC is 0.05, the rough design effect is 3.45. Dividing 500 by that penalty gives an effective sample size of about 145 for a mean-like quantity. That number is useful intuition, not a replacement for a design-specific power calculation. More importantly, it still does not turn 10 treatment assignments into 500.

Use the explorer below to see the distinction. Increasing the cluster size can add outcome information, but the number of clusters remains the number that carries the treatment allocation and the finite-sample degrees-of-freedom problem.

Interactive design-effect check

How many independent pieces of information do you really have?

Move the cluster count, average cluster size, and ICC. The rough effective sample size shows why adding patients inside the same sites cannot substitute for adding randomized or independent clusters.

ReadoutHigh risk of over-reading the patient count
Nominal patients500clusters × average size
Design effect3.45×rough correlation penalty
Rough effective n145not an inferential df
Clusters per arm*5.0if split evenly
Many patients may be repeating the same cluster-level information. Treat the nominal sample size as a weak reassurance, then focus on the number and allocation of clusters, the ICC assumptions, and finite-sample inference.

Teaching approximation for equal-sized clusters and a mean-like contrast. It is not a power calculation, does not estimate the ICC, and cannot decide whether a particular small-sample correction is adequate. *A cluster-level treatment effect is still informed by the number and allocation of clusters, not by this effective-n display.

The standard error can be “clustered” and still too optimistic

Researchers often write that they used cluster-robust standard errors and move on. That is an important correction when observations within clusters are correlated, but conventional robust variance estimators rely on large-sample behavior in the number of clusters. With few clusters, they can underestimate uncertainty. A model may then report a narrow confidence interval or a small p-value because the approximation treats the cluster count as larger than it is.

The same warning applies to a default normal or z-based test. A small number of clusters calls for an inferential method that respects finite-sample behavior: for example, a bias-corrected cluster-robust variance estimator with an appropriate t-reference distribution, a small-sample degrees-of-freedom correction for a mixed model, a cluster-level analysis when its assumptions fit the estimand, or a randomization-based procedure. The right choice depends on the design, outcome, treatment allocation, and estimand. “Use robust SEs” is not a complete method section.

Decision rule

When the treatment was assigned at the cluster level, read the cluster count before the patient count. If clusters are few or unevenly split between arms, demand a finite-sample inference strategy and an analysis that makes the cluster-level information visible.

Four ways a paper can look larger than it is

1. Counting people as degrees of freedom

A table highlights thousands of patients but buries the fact that only a small number of clinics received the intervention. The patient count is not the relevant denominator for cluster-level assignment.

2. Using a default z-test

The point estimate may be reasonable while the reference distribution is too optimistic. Small-cluster corrections usually widen uncertainty because the method has less independent information to work with.

3. Treating a random intercept as a cure

A mixed model can represent within-site correlation, but a random-effects label does not by itself guarantee calibrated confidence intervals with few clusters or a credible ICC estimate.

4. Ignoring arm imbalance

Ten total clusters are not equivalent to five versus five, three versus seven, or one versus nine. The number of clusters in each arm and the allocation mechanism matter for both power and inference.

There is no universal “enough clusters” number

Method papers often describe cluster randomized trials with fewer than roughly 30 to 40 clusters as small, but that is a warning benchmark, not a universal pass mark. The threshold changes with the number of arms, cluster-size variation, ICC, outcome distribution, covariate adjustment, effect size, and whether the intervention varies at the cluster or individual level. A balanced 28-cluster trial with a prespecified small-sample analysis is not the same inferential object as a 28-cluster observational study with one treated site and 27 controls.

Power calculations should therefore vary the number of clusters, cluster sizes, ICC, allocation ratio, and plausible outcome rates. If cluster sizes are unequal, the simple design-effect formula can be too reassuring. Large sites may add patients without adding proportional information, and cluster size can itself reflect referral intensity, staffing, or outcome capture. That is a separate estimand and weighting question, not merely a standard-error question.

What a defensible analysis reports

QuestionWhat to look forRed flag
What was randomized?The number of clusters, allocation by arm, and whether treatment varied at the cluster level.The abstract reports only individual participants.
How much information does each cluster add?Cluster-size distribution, ICC assumption or estimate, and a design-specific power calculation.A single ICC and average size with no sensitivity analysis.
Is inference calibrated for few clusters?Bias-corrected variance, degrees-of-freedom correction, randomization inference, or a justified alternative.Default robust SEs paired with a normal approximation.
What population does the effect represent?A clear individual-average or cluster-average estimand and an explanation of site weighting.“Adjusted for clustering” is treated as the entire design explanation.

A reviewer’s five-question sequence

  1. Count the clusters first. Record total clusters and clusters per arm before reading the participant denominator.
  2. Identify the assignment unit. If the intervention was allocated to clinics, wards, schools, or communities, the design lives at that level.
  3. Separate information loss from inference. Ask how the ICC and cluster sizes informed power, then separately ask how the final uncertainty was calibrated.
  4. Inspect imbalance and leverage. Look for one treated cluster, very large sites, cluster-level covariates, and sensitivity to leaving out a cluster.
  5. Read the estimand. Decide whether the result is an effect for the average person, the average cluster, or a policy delivered through clusters. Those are not interchangeable.

Where Aqrab fits

Small-cluster problems are exactly the sort of flaw that hides behind familiar method names. A paper can say “mixed effects,” “GEE,” or “cluster-robust” and still leave the key judgment unanswered: does the amount of independent cluster information support the precision being claimed?

If you are reviewing a cluster-randomized trial, stepped-wedge design, multicenter registry, or site-level intervention, try Aqrab for a methods critique pass. You can also use the developer workflows to put cluster count, assignment unit, estimand, and finite-sample inference on the review checklist before publication.

The practical bottom line

Adding patients inside a site can improve outcome measurement without creating another independent treatment assignment. With few clusters, the design may be informative but the default inferential shortcuts are not automatically trustworthy.

So ask the simplest question first: how many times did the study actually get to assign treatment? That number should shape the power calculation, the uncertainty method, and the confidence of the conclusion.

Further reading

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive