Four intervals for one proportion, and what each actually covers.
Four intervals, and what each actually covers
Wilson gives 5.24% to 36.04%. Wald gives 0% to 30.65% — a lower bound at zero it has no evidence for.
3 of 20 — observed proportion 15.0000%
95.000% Wilson interval: 5.2369% to 36.0419%
The Wald interval gives 0.0000% to 30.6491% — narrower than Wilson, and it covers the truth less often than it claims.
Four confidence-interval methods for the same proportion
Method
Lower
Upper
Width
Use it when
Wilson
5.2369%
36.0419%
30.8050pp
almost always — the default
Wald
0.0000%
30.6491%
30.6491pp
never, on this evidence
Agresti-Coull
4.3939%
36.8849%
32.4910pp
a simpler hand calculation
Clopper-Pearson (exact)
3.2071%
37.8927%
34.6856pp
when coverage must never fall short
Observed proportion
15.0000%
3 of 20
Wilson width
30.805pp
the recommended interval
Wald width
30.649pp
narrower than it has earned
Sample size check
Too small
for the normal approximation
What each method actually covers at n = 20
Exact, not simulated. For each true proportion, every possible outcome from 0 to 20 successes is enumerated and weighted by its binomial probability, so these are the real coverage rates rather than an estimate of them. The target is 95.000%.
Exact coverage of four interval methods across true proportions
True proportion
Wald
Wilson
Agresti-Coull
Clopper-Pearson
2.00%
33.18%
94.01%
99.29%
99.29%
5.00%
63.89%
92.45%
98.41%
98.41%
10.00%
87.60%
95.68%
95.68%
98.87%
20.00%
92.08%
95.63%
95.63%
97.85%
35.00%
93.60%
96.83%
96.83%
96.83%
50.00%
95.86%
95.86%
95.86%
95.86%
Amber marks coverage more than 2 points below target. Clopper-Pearson never falls short and is usually several points above, which is the price of its guarantee: it is wider than it needs to be.
Points are the Wald interval’s real coverage; the line across the middle is the level it claims. The gap between them is the error the textbook method makes, and it does not close at any sample size you are likely to have.
Wilson is derived by inverting the score test; Wald is derived by assuming the estimate is normal. That assumption is what fails. Wilson’s interval cannot leave [0, 1], cannot collapse to a point, and holds its nominal coverage far better at every sample size — and it was published in 1927, so the reason your textbook teaches the other one is not that it came first.
The Wilson bounds are exactly where the score test stops rejecting. Test each end as a hypothesised proportion and you get p = 0.050to six decimal places — the interval and the test are one procedure seen twice. That identity is checked on every bound from 1 to 39 successes out of 40 in the verification suite.
What this tool shows
A “95%” Wald interval covers the true proportion 80.85% of the time at n = 30, p = 0.1. Not estimated — enumerated, by summing every possible outcome weighted by its binomial probability. The tool computes four intervals on your counts and prints the real coverage of each at your sample size, so the choice between them stops being a matter of taste.
Wilson score interval, the default for almost every use
Wald, Agresti-Coull and Clopper-Pearson on the same counts
The exact coverage of each method at your n, across six true proportions
The degenerate cases: 0 of n and n of n, where Wald collapses to a point
Any confidence level, not just 95%
The identity linking the interval to the score test
Four methods Exact coverage Wald failure shown Any level
Updated 12 September 2026 · Works in any browser, no installation
Wilson’s interval is the set of proportions a score test would not reject — it inverts the test rather than assuming the estimate is normal. That single change fixes everything wrong with the textbook interval: it cannot leave [0, 1], it cannot collapse to a point, and it holds its stated confidence level far better at every sample size. It was published in 1927.
At a glance
Formula shown
Wilson: the interval is centred at (p̂ + z²/2n)/(1 + z²/n) with half-width z·√(p̂(1−p̂)/n + z²/4n²)/(1 + z²/n). The shift of the centre toward ½ and the extra z²/4n² term under the root are what the Wald formula omits. Wald: p̂ ± z·√(p̂(1−p̂)/n), which is zero-width when p̂ is 0 or 1. Agresti-Coull adds z²/2 successes and failures, then applies Wald. Clopper-Pearson inverts the binomial test exactly, using beta quantiles.
Scenario support
Conversion rates from an A/B test, defect rates in a batch, survey percentages, prevalence estimates, click-through rates, pass rates, and any “x out of n” figure that needs an uncertainty attached.
Educational estimate
Planning support from the values you enter — not professional advice.
What a “95%” interval actually covers
A 95% confidence interval means one thing: over repeated samples, it contains the true value 95% of the time. That is a claim which can be checked exactly, because for a binomial the entire sample space is enumerable.
The Wald interval does not deliver it. At n = 30 with a true proportion of 0.1, its real coverage is 80.85%. At n = 20, p = 0.1 it is 87.60%. At n = 100 with p = 0.02 — a large sample by most standards — it is 86.64%.
These are not simulations. For each true proportion, every possible outcome from 0 to n successes is weighted by its binomial probability and checked. There is no sampling error in the numbers above, and the tool recomputes them for your own n.
The shortfall does not shrink the way people expect. It improves with n for a fixed p, but it also depends on where p sits, and near 0 or 1 it persists at sample sizes in the hundreds. “np > 5” does not rescue it.
Wilson is closer to nominal at every grid point tested, and Clopper-Pearson never falls below it at all — coming in at 96.15% to 99.22% on the same grid, which is the other failure mode: an interval wider than the evidence requires.
So the ranking is not a matter of preference. Use Wilson unless coverage must be guaranteed never to fall short, in which case use Clopper-Pearson and accept the extra width.
Zero successes, and an interval of zero width
The clearest single argument against the Wald interval takes one line of arithmetic.
With 0 successes in 20 trials, p̂ is 0, so p̂(1 − p̂) is 0. The standard error is 0. The interval is 0 ± 0 — a “95% confidence interval” of [0, 0], asserting that the true proportion is exactly zero on the strength of twenty observations.
It is not zero. A proportion of 8% would produce zero successes in twenty trials about 19% of the time. Wilson gives [0, 0.1611], which is what the evidence actually supports; Clopper-Pearson gives [0, 0.1684].
The same collapse happens at the other end. Twenty successes in twenty trials gives Wald [1, 1]. Wilson gives a lower bound of 0.8389.
This is the case where the rule of three is usually reached for — the approximation that with no events in n trials, the upper 95% bound is about 3/n. At n = 20 that gives 0.15, close to Wilson’s 0.1611 and an order of magnitude away from Wald’s 0.
And zero events is common. No adverse reactions in a safety cohort, no failures in a burn-in test, no defects in a sampled batch: every one of these needs an upper bound, and the textbook formula returns the one number that cannot be true.
Why inverting the test works
The two formulas differ because they answer the question in opposite directions.
Wald starts from the estimate. It assumes p̂ is approximately normal around the true p, with a variance estimated from p̂ itself. Two approximations, and both are worst exactly where the interval is needed most: at small n and at p near a boundary.
Wilson starts from the hypothesis. For each candidate value of p, it asks whether a score test would reject it, using the variance that p implies rather than the one p̂ estimates. The interval is the set of candidates that survive.
That removes the second approximation entirely, because under each hypothesis the variance is known rather than estimated. The remaining normal approximation is applied to a statistic much closer to normal than p̂ is.
The algebra falls out as a shift and a widening. The centre moves from p̂ toward ½ by z²/2n, and an extra z²/4n² appears under the square root. Both vanish as n grows, which is why the two agree on large samples and why the difference is invisible in the examples textbooks use.
The identity is exact and the tool relies on it. Test either Wilson bound as a hypothesised proportion and the score test returns p = 0.05 to six decimal places. The verification suite checks that on every bound from 1 to 39 successes out of 40.
Choosing between the four
All four are in the tool because the comparison is the argument. Here is what each is for.
Wilson is the default. Best coverage of the three approximate methods across the whole range of p and n, no degenerate cases, and computable by hand. If one interval has to be chosen without knowing the data, this is it.
Agresti-Coull is Wilson made memorable. Add z²/2 successes and z²/2 failures — roughly two of each at 95% — then apply the Wald formula. It approximates Wilson’s behaviour with an easier rule, and is slightly more conservative.
Clopper-Pearson guarantees coverage and pays for it. It inverts the binomial test exactly, so coverage is never below nominal — and usually several points above, which means intervals wider than the data warrants. Use it where undercoverage is unacceptable: regulatory submissions, safety limits, anything audited.
Wald is here to be looked at. It is the interval in most textbooks and most spreadsheets, and the coverage table is the reason it should not be the one you report.
On large balanced samples they all agree, which is the honest caveat: with 512 of 1000, the four differ in the third decimal place. If your sample is large and your proportion is near ½, the choice does not matter. It matters everywhere else.
What a confidence interval does not say
The method is only half the problem. The sentence people write about the result is the other half.
It is not a 95% probability that the true value lies inside. The true proportion is a fixed number; it is either in this interval or it is not. The 95% describes the procedure — over many samples, 95% of the intervals it builds contain the truth.
It is not a range of plausible values for the next observation. That is a prediction interval, and it is much wider. A confidence interval is about the parameter, not about future data.
Two overlapping intervals do not mean no difference. Intervals for two proportions can overlap while a two-proportion test is significant. Comparing two groups needs an interval for the DIFFERENCE, not a visual check of two separate ones.
And it says nothing about bias. An interval accounts for sampling variation only. If the sample was not random with respect to the thing being measured, the interval will be narrow, confident and wrong — and no method in this tool can detect that.
Reporting a proportion with its interval
Four things, and the first is the one almost always missing.
Name the method. “12% (95% CI 8–17%)” does not say which of four methods produced those bounds, and on small samples they differ enough to matter. “95% Wilson CI” costs one word.
Give the raw counts. “3 of 20” carries the sample size and lets a reader recompute. “15%” carries neither, and 15% from 20 observations is a different finding from 15% from 20,000.
Do not report a proportion from a handful of observations without one. A bare percentage from n = 7 reads as precise and is not. If the interval is embarrassingly wide, that is the finding.
Round the bounds, not the estimate you compute from. Compute at full precision and round once at the end, and never round the interval so tightly that it no longer contains the point estimate.
Sources and methodology
References for interval estimation of a proportion.
Method. Coverage is computed EXACTLY rather than simulated: for a given true proportion, every outcome from 0 to n successes is weighted by its binomial probability and checked for containment, so the figures carry no sampling error. The suite runs that computation across a grid of sample sizes and proportions, confirming that Wald falls below nominal at every point — reaching 80.85% at n = 30, p = 0.1 — that Wilson is closer to nominal than Wald at every point, and that Clopper-Pearson never drops below it. The Wald degeneracy at 0 and n successes is asserted as an exact zero width rather than described. The identity between the Wilson bounds and the score test is verified on every bound from 1 to 39 successes out of 40. That engine is verified on every change against 75 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
One-Proportion Z-TestThe score z-test with the exact binomial test beside it: 60 of 100 against 0.5 gives p = 0.0455 by one and 0.0569 by the other — opposite verdicts at the conventional threshold, on identical data.
Two-Proportion Z-TestCompares two rates with the pooled z-test and a Newcombe interval for the difference, and shows that z² equals the 2×2 chi-square statistic exactly — so the “z-test or chi-square” question has no content.
Confidence IntervalIntervals for a mean or a proportion using t at every sample size and Wilson rather than the textbook Wald formula — with both methods shown, because Wald returns [0,0] at zero successes.
Margin of ErrorMargin of error for a percentage or an average, shown across seven sample sizes so the square-root law is visible — every doubling buys exactly 29.3%, never more.
Sample SizeResponses needed for a target margin of error, with the finite-population correction and a table of the whole cost curve — because n scales with 1/margin², so the last point of precision costs more than the first ten.
Binomial DistributionExact binomial probabilities at any n — including thousands, where a factorial overflows — with the normal approximation beside them and its error measured, which is 0.6% at the centre and 261% in the tail.
An educational tool. A confidence interval quantifies sampling variation only — it cannot account for selection bias, measurement error, or a sample that was not random with respect to what is being measured, and in those cases a narrow interval is misleading rather than reassuring.
Published four confidence intervals for one proportion with the EXACT coverage each delivers, enumerated rather than simulated: for a given true proportion every outcome from 0 to n successes is weighted by its binomial probability, so the figures carry no sampling error.
The headline is the textbook method failing. A "95%" Wald interval covers the true proportion 80.85% of the time at n = 30, p = 0.1; 87.60% at n = 20, p = 0.1; and 86.64% at n = 100, p = 0.02 - a large sample by most standards.
Clopper-Pearson never falls below nominal on the same grid, coming in at 96.15% to 99.22% - the other failure mode, an interval wider than the evidence requires.
The degenerate case is asserted as an exact zero: with 0 successes in 20 trials the Wald interval is [0, 0], a 95% confidence interval claiming certainty from twenty observations. Wilson gives [0, 0.1611].
The identity between the Wilson bounds and the score test is verified on every bound from 1 to 39 successes out of 40 - each returns p = 0.05 to six decimal places.
Add this calculator to your site
Responsive embed — and private: nothing your visitors type leaves their browser.