An effect size of 0.4 at 80% power needs 99 subjects per arm if individuals are randomised. Randomise clusters of thirty instead, at an intraclass correlation of 0.02, and the design effect is 1.58: 156 per arm, or 6 clusters of thirty. Note what the trade-off table shows — five per cluster would need 22 clusters and 106 subjects per arm, and a hundred per cluster would need 3 clusters and 293. Fewer clusters always means more people, and the exchange rate gets worse as the clusters grow.
d = 0.400 · ρ = 0.0200 · 30 per cluster · 80% power
6 clusters per arm — 12 clusters and 360 subjects in total
An individually randomised trial would need 99 per arm. The design effect of 1.5800 takes that to 156, which is 1.580 times the sample for the same power. With only 6 clusters per arm the between-cluster variance is estimated from very little, and the analysis will be unstable in a way this calculation does not capture.
Clusters per arm
6
below the usual floor of ten
Subjects per arm
156
99 if individually randomised
Design effect
1.5800
1.580× the sample
Total enrolment
360
across 12 clusters
Clusters against cluster size
Clusters and subjects needed per arm at a range of cluster sizes
Per cluster
Design effect
Clusters per arm
Subjects per arm
Against the smallest
5
1.080
22
106
1.00×
10
1.180
12
116
1.09×
15
1.280
9
126
1.19×
20
1.380
7
136
1.28×
30
1.580
6
156
1.47×
50
1.980
4
195
1.84×
100
2.980
3
293
2.76×
Reading down the table is the design decision. Clusters fall and subjects rise, and the exchange rate gets worse as the clusters grow — because effective observations per cluster approach a ceiling of 1/ρ and the extra people are buying less and less.
The clustered requirement is exactly the individual one multiplied by the design effect — there is no separate formula. The verification suite asserts that relation, along with the direction of the trade-off: larger clusters always need more total subjects and fewer clusters.
This arithmetic has no opinion about how few clusters is too few, and it should. With fewer than about ten clusters per arm the between-cluster variance is estimated from a handful of numbers, the degrees of freedom for the comparison collapse, and a small-sample correction becomes necessary — none of which shows up in a sample size that looks satisfied.
Clusters and subjects The whole trade-off Against individual randomisation Blind to too-few clusters
What this tool shows
At an effect size of 0.4 and an intraclass correlation of 0.02, five per cluster needs 22 clusters and 106 subjects per arm. A hundred per cluster needs 3 clusters and 293. Fewer clusters, nearly three times the people — and the exchange rate gets worse as the clusters grow, because each extra person is buying less. Recruiting sites is almost always the harder job, which is why designs drift towards fewer and larger clusters and end up underpowered.
Clusters per arm for a target power, and the subjects they contain
The individually randomised requirement for the same effect, so the cost of clustering is explicit
The full trade-off across seven cluster sizes, with clusters and subjects for each
The design effect that links the two, printed rather than buried
A warning when the design falls below the point where the between-cluster variance can be estimated
Why this arithmetic has no opinion about too-few clusters, and what that omits
Clusters and subjects The whole trade-off Against individual randomisation Blind to too-few clusters
Ten clusters per arm is a floor this arithmetic does not enforce.
Updated 13 September 2026 · Works in any browser, no installation
A cluster-randomised trial needs more subjects than an individually randomised one for the same power, and the multiplier is the design effect. Once that is applied, the remaining decision is how to split the requirement: many small clusters or few large ones. The arithmetic prefers many small ones consistently, because effective observations per cluster approach a ceiling while clusters themselves keep contributing in full.
At a glance
Formula shown
Per arm, an individually randomised trial needs n = 2(z_{α/2} + z_β)²/d². A clustered one needs that times the design effect 1 + (m − 1)ρ, and the number of clusters is the result divided by m. There is no separate clustered formula — the design effect is the whole of it, which is why the two requirements are printed side by side.
Scenario support
Designing a cluster-randomised trial in schools, clinics, villages or workplaces; deciding between recruiting more sites and enrolling more people at each; and checking whether a published cluster trial had enough clusters to estimate what it claimed.
Educational estimate
Planning support from the values you enter — not professional advice.
Clusters fall, subjects rise, and the rate gets worse
The table on this page is the design decision, and reading down it makes an argument that logistics usually resist.
Five per cluster: 22 clusters and 106 subjects per arm. The design effect is only 1.08.
Thirty per cluster: 6 clusters and 156 subjects. A design effect of 1.58.
A hundred per cluster: 3 clusters and 293 subjects. Nearly three times the people of the first row.
And the exchange rate worsens as you go down, because effective observations per cluster approach a ceiling of 1/ρ — the extra people are buying less and less, while each extra cluster keeps contributing in full.
The floor this calculation does not enforce
The sample size can be satisfied by a design that cannot actually be analysed, and nothing in the arithmetic says so.
Three clusters per arm satisfies the first preset at a hundred per cluster. The power calculation is content.
The between-cluster variance would be estimated from three numbers. Which is not an estimate in any useful sense.
The degrees of freedom for the comparison come from the clusters, not the subjects. With six clusters in total that is four, and a t distribution on four degrees of freedom is a very different thing from a normal.
Ten clusters per arm is the usual practical floor, below which a small-sample correction becomes necessary and the nominal power is optimistic. The page warns when a design falls under it, because the number itself will not.
What randomising groups costs
Every clustered design pays a price against individual randomisation, and the page prints both so the price is a number rather than an assumption.
The first preset: 99 per arm individually, 156 clustered. A 58% increase from a correlation of 0.02.
The second: 63 against 183. Nearly three times, from a correlation of 0.1 with twenty per cluster.
Sometimes the price is unavoidable. An intervention delivered to a whole school cannot be randomised pupil by pupil, and contamination between arms would be worse than the inflation.
Sometimes it is a convenience being paid for in sample. Worth knowing which, and the comparison is what makes the question answerable.
The correlation is the input you are least sure of
Effect size and power are choices. The intraclass correlation is a forecast, and the requirement is linear in it.
Plan with 0.01 when the truth is 0.02 and you have planned half the inflation. At fifty per cluster that is a design effect of 1.49 rather than 1.98.
Published correlations vary by setting and outcome. Clinical process measures cluster far more strongly than patient outcomes; school attainment clusters more than attendance.
A random-intercept model on existing data estimates it, unbiasedly and noisily — thirty clusters recovers the truth on average with a wide spread on any single estimate.
So plan at the upper end of a plausible range, and report the assumed value. A trial whose correlation turned out double what was planned is underpowered in a way that is only visible if the assumption was stated.
Unequal clusters cost more still
The formula assumes every cluster is the same size, and real ones are not.
Variation in cluster size adds to the design effect. The standard adjustment multiplies the correlation term by 1 plus the squared coefficient of variation of the sizes.
A coefficient of variation of 0.5 is common in practice — clinics and schools differ by more than people expect — and it inflates the correlation term by 25%.
This page uses the equal-size formula, so with highly variable clusters it understates the requirement and the page says so rather than implying otherwise.
The practical response is to plan on the average size and add a margin, or to compute the adjustment separately when the size distribution is known in advance.
The analysis has to match the design
A cluster-randomised trial analysed as though individuals were randomised produces standard errors that are too small, which is the same error the sample size was inflated to avoid.
The unit of randomisation sets the unit of analysis. Ignoring it in the analysis undoes the design effect the protocol paid for.
A mixed model with a random intercept for cluster is the standard approach, and it uses the actual cluster sizes rather than an assumed average.
With few clusters it needs a small-sample correction, because the degrees of freedom come from the clusters and the default approximation is optimistic.
A cluster-level summary analysis is the simple alternative: reduce each cluster to one number and run an ordinary comparison on those. Less efficient, and difficult to get wrong.
Reporting a cluster trial design
Four items, and the first is the one that most often goes missing from a protocol.
Give the number of clusters, not just the number of subjects. It is what most of the power rests on and what the degrees of freedom come from.
Give the assumed correlation and its source. The requirement is linear in it and nobody can check it otherwise.
Give the individually randomised comparison. It states the cost of the design choice in one number.
And say how cluster-size variation was handled. The equal-size formula understates the requirement when sizes vary, which they always do.
Method. The individually randomised requirement is the standard two-sample formula, and the clustered one is exactly that multiplied by the design effect — there is no separate expression, which is why both are printed and the multiplier between them is shown. The trade-off table recomputes the requirement at seven cluster sizes rather than interpolating, so each row is the answer the calculator would give for that design. The verification suite asserts the relation between the two requirements, and asserts the direction of the trade-off on the shipped curve: larger clusters always need more total subjects and fewer clusters. It also checks the boundary behaviour — a cluster of one and a zero effect size are both refused rather than returning a plausible number. That engine is verified on every change against 128 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
Design EffectThe design effect 1 + (m − 1)ρ with the effective sample size it implies, the ceiling of 1/ρ effective observations per cluster, and the marginal gain from each extra person.
Mixed ModelA random-intercept model from variance components, with the intraclass correlation, the design effect it implies, both standard errors for the overall mean, and shrunk group estimates.
Intraclass CorrelationAll six ICC forms from one subject-by-rater matrix, with the rater means that drive them apart: one 8x3 matrix gives ICC(1,1) = 0.1277 and ICC(3,1) = 0.9852.
Sample SizeResponses needed for a target margin of error, with the finite-population correction and a table of the whole cost curve — because n scales with 1/margin², so the last point of precision costs more than the first ten.
Survival Sample SizeSchoenfeld's event requirement converted to an enrolment target, with the sensitivity to the hazard ratio and the power achieved at every event count.
Statistical PowerPower and sample size from the non-central t rather than a normal approximation, with the gap shown — plus a live demonstration that post-hoc power is a function of the p-value alone, and 0.500044 at p = 0.05 for every study ever run.
An educational tool. This uses the equal-cluster-size formula, which understates the requirement when cluster sizes vary — and they always do. The calculation is also silent about having too few clusters: a design with three per arm can satisfy the sample size while leaving the between-cluster variance estimated from three numbers and the comparison on very few degrees of freedom.