Power before the data, from the right distribution — and why post-hoc power says nothing.
Power, and the sample size that buys it
The textbook 80% design: Cohen’s medium effect at 64 per group.
d = 0.500, 64 per group, α = 0.050
Power = 80.15%
At or above the conventional 80% floor.
Power
80.15%
Type II error β
19.85%
misses a real effect
Non-centrality δ
2.8284
d·√(n/2)
Total observations
128
126 degrees of freedom
Normal approximation
80.74%
what most tools print
Overstatement
0.60 pts
optimistic
Critical t
1.9790
two-sided, 126 df
n for 80% power
64
per group
The normal approximation says 80.74%; the non-central t says 80.15%. Under a real alternative the t statistic follows a non-central t, not a normal. Substituting a normal ignores that the pooled standard deviation is estimated rather than known, and the error runs in one direction: optimistic, and worst at small n — which is precisely when a power calculation is being done to justify a small n.
observed power = 41.722%
Notice that no sample size and no effect size were needed to compute that. Observed power is a strictly decreasing function of the p-value and nothing else. At p = 0.05 it is always 0.500044 — exactly one half plus Φ(−2z), for any study ever run — and it crosses one half at p = 0.050013. So “the result was non-significant and post-hoc power was low” restates the p-value in different units. It cannot tell you whether the effect was real, and it makes the study that came closest to significance look the best powered. Power has to be computedbefore the data, at an effect size you would care about — which is what the panel above is for.
80% is a convention, not a standard. Cohen proposed it as a reasonable default when nothing else is known, on the reasoning that a Type I error is about four times as costly as a Type II. If that ratio is wrong for your question — a screening study where a miss is expensive, a confirmatory trial where a false positive is — then 80% is the wrong number, and the honest move is to say which ratio you chose rather than to inherit his.
What this tool shows
Post-hoc power is the p-value wearing different units. At p = 0.05 observed power is 0.500044 — for every study ever run, whatever its sample size and whatever its effect. The tool computes it from a p-value and nothing else, so you can watch that happen. Power itself uses the non-central t, which at n = 10 gives 18.5% where the usual approximation claims 20.1%.
Power from the non-central t, with the normal approximation beside it
Sample size for a target power, by search rather than by a formula that rounds short
One-sample, two-sample and paired designs
A power curve against n, with the 80% line drawn
Observed power from a p-value — computed in order to argue against it
Cohen’s h, so a 5-point gap at 2% is not treated like one at 50%
Non-central t Post-hoc power, refuted Power curve Three designs
Updated 12 September 2026 · Works in any browser, no installation
Power is the probability of detecting an effect that is really there. It depends on the effect size, the sample size and the significance level — and it is a property of a design, computed before the data, which is what separates it from the number people compute afterwards and call by the same name.
At a glance
Formula shown
Under the alternative the t statistic follows a NON-CENTRAL t with non-centrality δ = d·√(n/2) for two groups of n, or d·√n for one sample or a paired design, on the same degrees of freedom as the test. Power is the probability that |T| exceeds the central critical value, computed from Lenth’s AS 243 series. Sample size is found by searching n rather than from n = (z₁₋α/₂ + z_power)²·k/d², which uses normal quantiles where the test uses t ones and lands short.
Scenario support
Planning a study before collecting data, sizing an A/B test, checking whether a published null result had any chance of detecting the effect it was looking for, and deciding whether a proposed sample is worth running at all.
Educational estimate
Planning support from the values you enter — not professional advice.
Post-hoc power is the p-value in different units
The request arrives from reviewers constantly: the result was not significant, so report the achieved power. It is a request for a number that cannot answer the question.
Observed power is computed from the effect the study found, which is the same quantity the p-value was computed from. Put the algebra together and the sample size cancels: observed power is Φ(zobs − zcrit) + Φ(−zobs − zcrit), and zobs is just the p-value restated.
So it is the same number twice. At p = 0.05 observed power is 0.500044 — exactly one half plus Φ(−2zcrit) — for a study of ten people or ten thousand. At p = 0.10 it is 0.376; at p = 0.20, 0.249. The tool takes a p-value and nothing else, which is the demonstration.
Which makes the usual sentence a tautology. “Non-significant, and post-hoc power was low” is guaranteed: every result with p above 0.050013 has observed power below one half, by arithmetic. It is not evidence that the study was too small. It is the definition of a p-value above 0.05.
And it inverts the ranking. Hoenig and Heisey call this the power approach paradox: among two non-significant studies, the one with the smaller p-value — the one closer to finding something — scores the higher observed power, and is therefore declared the better designed. Observed power rewards getting lucky.
Power has to be computed before the data exists. At an effect size you would care about, chosen because it is meaningful rather than because it is what you happened to observe. That is the panel above this one, and it is the only power calculation that answers anything.
The non-central t, and what the approximation costs
Under the null the t statistic follows a t distribution. Under a real alternative it does not — it follows a non-central t, and the difference is not cosmetic at the sample sizes people actually ask about.
Most online power calculators substitute a normal. It has a closed form, it is easy to invert for sample size, and at n = 200 it is right to three decimal places. At n = 10 it is not.
The error runs one way. The normal ignores that the standard deviation is estimated rather than known, so it understates the variability of the statistic and overstates the chance of clearing the threshold. At d = 0.5 with ten per group it claims 20.1% power against a true 18.5%.
Which is exactly the wrong direction. A power calculation at n = 10 is almost always being done to justify running a small study. An optimistic answer there is worse than no answer.
The sample-size formula inherits the same problem. n = (z1−α/2 + zpower)²·k/d² appears in every methods chapter and uses normal quantiles for a test that will use t ones. At d = 0.5 it returns 63 per group where 64 is needed; at d = 0.8, 25 where 26 is. One short, reliably. This tool searches n against the true power instead, and prints what the formula would have said.
The implementation is checked against simulation, not against another formula. Two hundred thousand simulated t tests per case, agreeing within a couple of standard errors, plus the requirement that power at a zero effect comes out at exactly α — which is the one value that has a known answer rather than a computed one.
Halving the effect quadruples the sample
The single most useful thing a power calculation tells you is not a number but a shape, and it is the reason study budgets behave the way they do.
n scales as 1/d². Detecting an effect half as large takes about four times as many people; a quarter as large, sixteen times. At 80% power and the 5% level: d = 0.8 needs 26 per group, d = 0.4 needs 100, d = 0.2 needs 394.
So an argument about whether an effect is “small” is an argument about budget. Deciding you would still care about half the effect you first proposed has just quadrupled the cost of the study, and that consequence is usually invisible in the discussion where the decision is made.
It also explains why underpowered studies cluster. Most real effects in most fields are small. A field where d = 0.2 is typical needs 788 people for a single two-group comparison, and most published studies in such fields have fewer than 100 — so most of them have power around 15%, and the significant results that emerge are systematically the ones where noise helped.
Which is why an underpowered significant result overstates the effect. If only large observed effects clear the threshold, then every published effect from a low-power design is biased upward. The fix is not a bigger p-value threshold; it is a bigger n, and the effect size reported alongside.
Pairing is the cheapest way to buy power. A paired design tests the differences, so it removes between-subject variability entirely. If the pairing is good, the same substantive effect becomes a much larger d, and the sample needed falls by the square of that improvement — which the paired-test framing makes concrete.
Where 80% came from, and when it is wrong
Eighty percent is quoted as though it were a standard. It is a default that one person proposed, with a stated reason, and the reason is checkable against your own situation.
Cohen suggested it because of a ratio. If a false positive is about four times as costly as a false negative, then β = 4α is the balanced choice, and at α = 0.05 that is β = 0.20 — power 0.80. It was explicitly offered as a convention for use when nothing better is known.
The ratio is often wrong. In a screening study a miss sends a real signal to the bin; in a confirmatory trial a false positive puts a useless treatment into practice. Those are not the same 4:1 trade, and inheriting Cohen’s number without checking is inheriting an assumption about costs you never made.
90% and 95% are routine where a miss is expensive, and the price is steep: going from 80% to 90% at a fixed effect size costs about 34% more sample, and to 95% about 66% more. The tool will compute either.
α is equally negotiable and rarely negotiated. Loosening it raises power for free in sample-size terms, and raises false positives. Tightening it does the reverse. Both numbers encode the same trade-off from opposite ends, and pretending only one of them is a choice is how 0.05 became a fact.
The honest reporting is to state both and why. “Powered at 90% to detect d = 0.4 at α = 0.01, because a false positive here means a recall” is a design. “Powered at 80%” is a habit.
Proportions are not differences
Sizing an A/B test on “we want to detect a 5-point lift” is the most common way a power calculation goes wrong, and the reason is that a proportion’s variance depends on the proportion.
Five points at 50% and five points at 2% are different problems. A rate near one half has maximum variance; a rate near zero has very little. The same absolute gap is a far larger effect at the low end, and a far easier one to detect per observation.
Cohen’s h measures it correctly. h = 2·arcsin√p₁ − 2·arcsin√p₂ — the arcsine transform is the one that makes a proportion’s variance constant. For 0.50 against 0.55 it gives 0.1002; for 0.02 against 0.07 it gives 0.2517, a 2.51 times larger effect from the same five points.
Which flips the intuition about which test is cheaper. Detecting a five-point lift on a 2% conversion rate needs roughly a sixth of the sample that the same lift needs at 50% — so the low-baseline test that feels hopeless is often the affordable one, provided the lift really is absolute rather than relative.
Relative and absolute lifts are also routinely confused. “A 10% improvement” on a 2% baseline is 2.2%, not 12%, and the two require sample sizes differing by orders of magnitude. Say which one you mean before computing anything.
And a power calculation is not a stopping rule. Checking an A/B test repeatedly and stopping when it turns significant inflates the false positive rate far past α, regardless of how the sample was sized. Sequential designs exist for that; the fix is not to peek more carefully.
What to report after a null result instead
If post-hoc power is the wrong answer to “was this study big enough”, something has to be the right one. Three things are.
The confidence interval, first. It contains everything observed power was reaching for and more: a narrow interval around zero says the effect is small; a wide one says the study could not tell. Same data, same information, no paradox — and the interval is already in your output.
The effect size with its interval, second. “d = 0.12, 95% CI −0.15 to 0.39” says precisely what was learned. It rules out large effects in both directions and admits that small ones remain possible, which is usually the honest summary of a null result.
An equivalence test, third, if the question was really about absence. Failing to reject is not evidence of no effect. Two one-sided tests against a pre-specified bound answer “is the effect smaller than I would care about”, which is the question people usually meant.
A power calculation at a meaningful effect is still worth reporting — computed before the data, at the effect you designed to detect, not at the one you found. “This study had 80% power to detect d = 0.5, and observed d = 0.08” is informative. Recomputing power at d = 0.08 is not.
None of this rescues a study that was too small. It does describe accurately what was learned from it, which is the most that can be asked after the fact — and it avoids the circularity of using the result to grade the design that produced it.
Sources and methodology
References for the method and for the post-hoc power argument.
Method. Power comes from the non-central t via Lenth’s AS 243 series, with the non-centrality d·√(n/2) for two independent groups and d·√n for one-sample and paired designs. The normal approximation is computed alongside rather than instead, so the gap is visible. Sample size is found by bracketing and bisecting on the true power rather than from the closed form, because power is a step function of an integer n and the published formula uses normal quantiles for a test that will use t ones. The suite checks the implementation against a 200,000-trial simulation of the actual t test at four designs, requires that power at a zero effect equals α exactly, confirms the non-central t at δ = 0 reproduces the central t, and pins the textbook sample sizes — 64, 394 and 26 per group at d = 0.5, 0.2 and 0.8. It also asserts the identity this page rests on: that observed power is strictly decreasing in p across 998 values and equals 0.5 + Φ(−2z) at p = α. That engine is verified on every change against 59 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
Sample SizeResponses needed for a target margin of error, with the finite-population correction and a table of the whole cost curve — because n scales with 1/margin², so the last point of precision costs more than the first ten.
Effect SizeCohen d, Hedges g and the overlap between groups, with a sample-size control that moves the p-value while leaving the effect size fixed — the same d gives t = 1.29 at n=30 and 23.57 at n=10,000.
Confidence IntervalIntervals for a mean or a proportion using t at every sample size and Wilson rather than the textbook Wald formula — with both methods shown, because Wald returns [0,0] at zero successes.
p-valueA p-value from a t or z statistic, one- or two-tailed — with a panel that holds an effect fixed and grows the sample, so you can watch significance appear from nothing but n.
t-testOne-sample, two-sample and paired t-tests defaulting to Welch, with Student's pooled version printed beside it — and a warning when the two disagree on the verdict.
Margin of ErrorMargin of error for a percentage or an average, shown across seven sample sizes so the square-root law is visible — every doubling buys exactly 29.3%, never more.
An educational tool. Power depends on assumptions about the effect size and the variance that are made before the data exists; observed or post-hoc power computed from a result carries no information beyond that result’s p-value and should not be reported as evidence about a design.
Published a power calculator that uses the NON-CENTRAL t rather than the normal approximation almost every other online tool substitutes. Under a real alternative the t statistic follows a non-central t, and the normal ignores that the standard deviation is estimated — at d = 0.5 with ten per group it claims 20.1% power against a true 18.5%, and the error runs optimistic at exactly the sample sizes a power calculation is done to justify.
Computes observed (post-hoc) power from a p-value and NOTHING ELSE, in order to argue against reporting it. The sample size and the effect size cancel out of the algebra: at p = 0.05 observed power is 0.5000443 — exactly one half plus Phi(-2z) — for every study ever run, and it crosses one half at p = 0.050013. Hoenig and Heisey's power approach paradox follows: among non-significant studies the one closest to significance scores the highest observed power.
Finds sample size by searching n against the true power rather than from n = (z + z)^2 k/d^2, which uses normal quantiles for a test that will use t ones and lands one short — 63 per group where 64 is needed at d = 0.5, and 25 where 26 is needed at d = 0.8. The closed-form answer is printed alongside so the shortfall is visible.
Reports Cohen's h for proportions rather than the raw difference, because a proportion's variance depends on the proportion: the same five-point gap is a 2.513 times larger effect at a 2% baseline than at 50%. Verified against a 200,000-trial simulation of the actual t test at four designs, and against the requirement that power at a zero effect equals alpha exactly.
Add this calculator to your site
Responsive embed — and private: nothing your visitors type leaves their browser.