Math calculator

Cohen's h Calculator

The effect size for two proportions.

Two proportions, on a fair scale

A doubling of a 1% conversion rate. h = 0.083459 and 1,127 per group for 80% power — the smallest sample of any preset here, for the smallest absolute gain.

φ₁ = 0.283794, φ₂ = 0.200335

h = 0.083459

A negligible effect by Cohen's bands, from a raw difference of 1.000 percentage points. Detecting it with 80% power at α = 0.05 two-sided needs 1,127 per group.

Cohen's h

0.083459

negligible

Raw difference

1.000 pp

what h is NOT

n per group

1,127

80% power, α = 0.05

Ratio

2.0000×

which h is also not

One percentage point, at nine places on the scale

Every row is the same absolute gain. Only where it sits changes.

Cohen’s h and the sample size it implies for a one-percentage-point gap at nine baselines
ComparisonRaw gapCohen’s hn per group for 80% power
1% → 2%1.00 pp0.0834591,127
5% → 6%1.00 pp0.0439074,072
10% → 11%1.00 pp0.0326297,373
20% → 21%1.00 pp0.02477212,791
30% → 31%1.00 pp0.02172116,637
50% → 51%1.00 pp0.02000119,620
70% → 71%1.00 pp0.02192816,323
90% → 91%1.00 pp0.0341166,744
98% → 99%1.00 pp0.0834591,127

The same one-point improvement needs 1,127 per group at a 1% baseline and 19,620 at 50% — seventeen times as much data for an identical absolute gain. That is not a quirk of the effect size; it is a real property of binomial variance, which is largest at 0.5.

h = 2·arcsin√p₁ − 2·arcsin√p₂ Bands 0.2 / 0.5 / 0.8, and they are conventions The sample size does not depend on the proportions, only on h

What this tool shows

Improving 1% to 2% gives h = 0.083459 and needs 1,127 per group for 80% power. Improving 50% to 51% gives h = 0.020001 and needs 19,620. Identical one-percentage-point gains, seventeen times the data. That is not a quirk of the effect size — binomial variance is largest at 0.5 and smallest at the ends, and the arcsine transform is what makes an effect size respect it.

  • Cohen’s h, with both arcsine-transformed proportions shown
  • The per-group sample size for 80% power at α = 0.05, two-sided
  • The same one-percentage-point gap at nine points on the scale
  • Cohen’s 0.2 / 0.5 / 0.8 bands, labelled as the conventions they are
  • The raw difference and the ratio beside h, so the three are not confused
  • Proportions of exactly 0 and 1 handled rather than returning infinity
Two proportions Arcsine transform Sample size Scale-position sweep

h is not a percentage-point difference and not a ratio.

Updated 12 September 2026 · Works in any browser, no installation

Cohen’s h is the difference between two proportions after each has been arcsine-square-root transformed, which puts them on a scale where a unit of difference costs the same amount of data everywhere. It is the proportion equivalent of Cohen’s d, and it exists because a raw percentage-point difference does not have that property: a gap near 0 or 1 is far easier to detect than the same gap at 0.5.

At a glance

Formula shown
h = 2·arcsin(√p₁) − 2·arcsin(√p₂). The transform φ = 2·arcsin(√p) is variance-stabilising for a binomial proportion: the variance of φ̂ is approximately 1/n regardless of p, which is exactly why the sample size formula n = (z_{α/2} + z_β)²/h² contains no p at all. With 80% power at α = 0.05 that constant is (1.959964 + 0.841621)² = 7.8489.
Scenario support
A/B testing and conversion-rate experiments, planning the sample size for a two-arm trial with a binary endpoint, comparing response rates between groups, survey question comparisons, and any power calculation where the outcome is a proportion rather than a mean.
Educational estimate
Planning support from the values you enter — not professional advice.

Where on the scale a gap sits changes what it costs

The tool’s sweep holds the raw difference at exactly one percentage point and moves it along the scale. Nothing else changes, and the sample size moves by a factor of seventeen.

1% to 2%: h = 0.083459, 1,127 per group. 5% to 6%: 0.043907 and 4,072. 10% to 11%: 0.032629 and 7,373. 30% to 31%: 0.021721 and 16,637. 50% to 51%: 0.020001 and 19,620.

Then it comes back down. 90% to 91% gives 0.034116 and 6,744; 98% to 99% gives 0.083459 and 1,127 — exactly the same as 1% to 2%, because the transform is symmetric.

The cause is binomial variance, which is p(1−p) and peaks at 0.5. Near the ends the data is less noisy, so a small difference stands out; in the middle it is buried.

Which is why a conversion test on a 1% page finishes far sooner than one on a 50% page, for the same absolute improvement — a fact that surprises people planning experiments and is entirely predictable from this table.

What the arcsine transform is actually for

It looks like an arbitrary bit of trigonometry dropped into a definition. It is the specific function that makes the variance stop depending on p.

The variance of a sample proportion is p(1−p)/n, which changes by a factor of 25 between p = 0.5 and p = 0.01. Any effect size built on the raw difference inherits that.

The variance of 2·arcsin(√p̂) is approximately 1/n, whatever p is. That is what “variance-stabilising” means, and it is why the transform is not interchangeable with a logit or any other reshaping.

The payoff is a sample-size formula with no p in it. n = (zα/2 + zβ)²/h² per group — 7.8489/h² for 80% power at α = 0.05 — which is the whole reason h is the effect size used for planning proportion studies.

It also means h is not interpretable as a difference in anything real. It is on the transformed scale, and the only honest readings of it are the conventional bands and the sample size it implies.

h, the risk difference and the risk ratio measure different things

Three summaries of the same two proportions, and choosing between them is a decision about what question you are answering.

The raw difference is what a person experiences. One extra percentage point is one extra person in a hundred, wherever on the scale it sits, and for a decision about individuals it is usually the right number — see number needed to treat.

The ratio is what a headline reports. 1% to 2% is a doubling and 50% to 51% is a 2% increase, which is why the first sounds impressive and the second does not.

h is what a power calculation needs. It is the only one of the three that maps directly onto how much data is required, which is a planning question rather than a reporting one.

They disagree in both directions. 1% to 2% has the smallest raw difference of the three framings and the largest ratio and the largest h of any one-point gap. Quoting one and implying another is the commonest way these get misused.

The 0.2 / 0.5 / 0.8 bands are conventions Cohen apologised for

Cohen offered them as a fallback for researchers with no basis for judging a meaningful effect, and was explicit that they were a last resort rather than a standard.

They carry no information about your setting. A “small” h of 0.083 between two vaccine arms can be enormously consequential; a “large” h of 0.8 on a marketing headline may be worth nothing.

The bands are also not comparable across effect-size families. h = 0.5 and d = 0.5 are both “medium” and are not equivalent quantities.

The useful question is the minimum effect worth detecting, which comes from the cost of the intervention and the value of the outcome — not from a table.

The tool prints the band because it is asked for, and prints the sample size beside it because that is the number the band is a proxy for.

The sample size here is a floor, not a plan

n = 7.8489/h² per group is the standard two-sided calculation at 80% power, and there are four reasons a real study needs more.

It assumes equal group sizes. An unbalanced allocation needs a larger total — a 1:2 split needs about 12.5% more people than 1:1 for the same power.

80% power means a one-in-five chance of missing a real effect. That is a conventional tolerance, not a safe one, and 90% power costs about a third more.

It uses the normal approximation. With very small expected counts a continuity correction or an exact test is more appropriate, and both need slightly more data.

And it assumes h is known. Powering a study on an effect size estimated from a small pilot is powering it on a number that is probably too large, since the pilots that get followed up are the ones that overshot.

Attrition is on top of all of it. The calculation gives analysable observations, not people recruited.

Reporting Cohen's h

Three things, and the first is the one that makes h interpretable at all.

Always give both raw proportions. h alone is on a transformed scale that nobody reads intuitively, and the two percentages take four characters.

Give h when the point is planning, and the risk difference when the point is consequence. They answer different questions and the audience for each is different.

State the power and alpha behind any sample size you quote. “1,127 per group” means nothing without “80% power, α = 0.05 two-sided” attached.

And if you use Cohen’s bands, say so. “A small effect by Cohen’s convention” is checkable; “a small effect” is an opinion.

Sources and methodology

References for Cohen's h and the arcsine transform.

Method. h is computed from the arcsine transform directly rather than from an approximation, and the sample size uses exact normal quantiles rather than the rounded 1.96 and 0.84 — which is why the constant is 7.8489 and the reported figures are 1,127 and 19,620 rather than the 1,126 and 19,598 the rounded constants give. The scale-position table recomputes h at each baseline rather than interpolating, which makes the seventeen-fold sample-size swing a measurement of the arcsine scale rather than an illustration of it. The symmetry h(p, q) = −h(q, p) and h(1−p, 1−q) = −h(p, q) are both asserted numerically across generated pairs. Proportions of exactly 0 and 1 transform to 0 and π rather than diverging, so they are handled rather than excluded. That engine is verified on every change against 115 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

Effect SizeCohen d, Hedges g and the overlap between groups, with a sample-size control that moves the p-value while leaving the effect size fixed — the same d gives t = 1.29 at n=30 and 23.57 at n=10,000.
Sample SizeResponses needed for a target margin of error, with the finite-population correction and a table of the whole cost curve — because n scales with 1/margin², so the last point of precision costs more than the first ten.
Two-Proportion Z-TestCompares two rates with the pooled z-test and a Newcombe interval for the difference, and shows that z² equals the 2×2 chi-square statistic exactly — so the “z-test or chi-square” question has no content.
Statistical PowerPower and sample size from the non-central t rather than a normal approximation, with the gap shown — plus a live demonstration that post-hoc power is a function of the p-value alone, and 0.500044 at p = 0.05 for every study ever run.
Number Needed to TreatNNT from the absolute risk reduction, with the relative figure beside it — two trials reporting the identical “50% reduction” have NNTs of 7 and 1,000, and the common shortcut says 2 for both.
Wilson Score IntervalWilson, Wald, Agresti-Coull and Clopper-Pearson on one set of counts, with the EXACT coverage each delivers at your sample size — a "95%" Wald interval covers the truth 80.85% of the time at n = 30, p = 0.1.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool. Cohen’s h lives on an arcsine-transformed scale and is not a percentage-point difference or a ratio, and the sample size it implies is a floor for an ideal balanced two-arm comparison — real studies need more for unequal allocation, attrition, and the optimism in any effect size estimated from a pilot.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (5 updates)

Published 12 September 2026

  1. Published the effect size for two proportions with the sample size it implies printed beside it, since the sample size is the number the effect size is a proxy for.
  2. Measured what the arcsine transform is actually doing: the same one-percentage-point gain gives h = 0.083459 at a 1% baseline and 0.020001 at 50%, which is 1,127 per group against 19,620 for 80% power. Seventeen times the data for an identical absolute improvement.
  3. Showed the symmetry as well as the peak — 98% to 99% needs exactly the same 1,127 as 1% to 2% — so the effect is the scale rather than anything about rarity.
  4. Used exact normal quantiles rather than the rounded 1.96 and 0.84, which is the difference between 1,127 and 1,126 at one end and 19,620 and 19,598 at the other.
  5. Verified that n per group is exactly 7.8489/h^2 with no dependence on p at all, which is the entire point of a variance-stabilising transform, and that h(1-p,1-q) = -h(p,q) to machine precision.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.