d is 0.6251 and g is 0.6013, a 3.81% correction. The interval is −0.2253 to 1.4279. A “medium to large” effect whose interval spans from a moderate harm to a very large benefit is the normal state of a twenty-two person study, and the point estimate on its own conceals it.
df = 20 · J = 0.961945
g = 0.6013 (95% CI -0.2253 to 1.4279)
Cohen's d on the same data is 0.6251, so the small-sample correction removes 3.81% of it. The pooled standard deviation is 11.9982. The interval contains zero: this design cannot distinguish the effect from none.
Hedges' g
0.6013
SE 0.4217
Cohen's d
0.6251
3.81% larger than g
Exact J
0.961945
shortcut gives 0.962025
Pooled SD
11.9982
df = 20
The variance of d is (n₁+n₂)/(n₁n₂) + d²/(2(n₁+n₂)), which here is 0.19221. The first term depends on the allocation, not only the total: at 12 and 10 it is 0.18333, against 0.18182 for the same people split evenly.
Exact against approximate correction
The exact gamma-function correction against the textbook shortcut at several degrees of freedom
df
Exact J
1 − 3/(4df − 1)
Error
d overstated by
4
0.79788456
0.80000000
2.115e-3
20.21%
6
0.86862669
0.86956522
9.385e-4
13.14%
10
0.92274561
0.92307692
3.313e-4
7.73%
20
0.96194453
0.96202532
8.078e-5
3.81%
38
0.98011040
0.98013245
2.205e-5
1.99%
100
0.99247805
0.99248120
3.148e-6
0.75%
200
0.99624452
0.99624531
7.842e-7
0.38%
398
0.99811420
0.99811439
1.977e-7
0.19%
At df = 4 the exact factor is √(2/π) = 0.79788456 in closed form, and d overstates the population effect by 20.21%. The shortcut’s error falls fourfold every time df doubles, which is why it is harmless above about forty and worth avoiding below ten.
Exact gamma correction d printed beside it Interval, not just a point Corrects bias, not confounding
What this tool shows
At four per group, Cohen’s d overstates the population effect by 13.14% — 1.2095 against a corrected 1.0506. This page applies the exact correction factor from the gamma function rather than the 1 − 3/(4df − 1) shortcut, and prints both so the difference is visible. It also shows what the shortcut hides: sixty-six participants split 33/33 and 60/6 give an identical g and an interval 71.2% wider.
Hedges’ g with the exact correction J = Γ(df/2)/(√(df/2)·Γ((df−1)/2))
Cohen’s d on the same data, and the percentage the correction removes
A confidence interval on g, which most effect-size calculators omit entirely
A table of exact against approximate J across eight degrees of freedom
The allocation effect: the same total n split evenly or lopsidedly
The pooled standard deviation and the variance of d, printed rather than hidden
Exact J d beside it Interval included Allocation shown
This corrects a small-sample bias. It cannot correct a biased study.
Updated 13 September 2026 · Works in any browser, no installation
Hedges’ g is Cohen’s d with the small-sample bias taken out. The pooled standard deviation in the denominator of d is a biased estimate of the population SD, and it is biased downward — so d comes out too large, systematically, and worst when the sample is small. Multiplying by a correction factor that depends only on the degrees of freedom fixes it. Above about fifty per group the correction is cosmetic; below ten it is not.
At a glance
Formula shown
d = (M₁ − M₂)/s_pooled with s_pooled = √(((n₁−1)s₁² + (n₂−1)s₂²)/(n₁+n₂−2)); g = J·d with the exact J = Γ(df/2)/(√(df/2)·Γ((df−1)/2)) on df = n₁+n₂−2. The variance is Var(d) = (n₁+n₂)/(n₁n₂) + d²/(2(n₁+n₂)) and Var(g) = J²·Var(d). The first term of that variance is where the allocation shows up: for a fixed total it is smallest when the two groups are equal and grows without limit as the split becomes lopsided.
Scenario support
Reporting an effect size from a small trial or pilot, preparing estimates for a meta-analysis where g is the standard input, comparing results across studies on different measurement scales, power planning where an uncorrected pilot d would inflate the assumed effect, and any write-up that needs an interval rather than a bare effect-size number.
Educational estimate
Planning support from the values you enter — not professional advice.
How much d overstates, at every sample size
The bias in d is not random noise. It is a systematic overstatement with a known size, and the table on this page prints it.
At df = 4 — three per group — the correction factor is exactly √(2/π) = 0.79788456. That is a closed form, not an approximation, and it means d is 20.21% too large.
At four per group (df = 6) the factor is 0.868627 and d overstates by 13.14%. The shipped pilot preset shows it: d = 1.2095, g = 1.0506.
At 12 versus 10 it is 3.81%. At 200 per group it is 0.19%. Which is why the correction is worth applying on a pilot and irrelevant on a large trial.
The direction never changes: g is always closer to zero than d. Any calculator that returns a g larger in magnitude than its own d has made an error, and the verification suite for this page asserts the inequality on 200 generated designs.
The exact factor, not the textbook shortcut
Nearly every implementation uses J ≈ 1 − 3/(4df − 1). It is a good approximation and it is not the definition, so this page computes the definition and shows the gap.
The exact factor is a ratio of gamma functions, evaluated here through log-gamma to avoid overflow at large df.
The shortcut errs by 9.385e-4 at df = 6 and 1.977e-7 at df = 398. Both are in the table, along with six intermediate values.
The error falls fourfold every time df doubles — it is O(1/df²) — which the suite asserts across five doublings rather than stating as a claim.
So the shortcut is harmless above about forty degrees of freedom and worth avoiding below ten, which is precisely where anyone is reaching for g in the first place.
Same 66 people, 71% wider interval
Effect-size calculators almost always ask for the total sample and stop there. The two balanced and unbalanced presets on this page show why that is not enough.
Sixty-six participants, the same means, the same standard deviations. Split 33/33 and split 60/6.
d = 0.6250, g = 0.6176, df = 64 and J = 0.988228 in both. Every headline number is identical.
The intervals are 0.1293 to 1.1060 and −0.2184 to 1.4536. The unbalanced one is 71.2% wider and it crosses zero.
The cause is the (n₁+n₂)/(n₁n₂) term in the variance, which is 0.0606 at 33/33 and 0.1833 at 60/6. Push it to 62/4 and the interval doubles. A lopsided design wastes participants, and the effect size alone will never tell you.
An effect size without an interval is half a result
The most common way to misuse g is to report it alone and then compare it to a threshold.
The pilot preset gives g = 1.0506 — a “large” effect by any convention. Its interval runs from −0.2587 to 2.3599.
That interval contains no effect, a trivial effect and an enormous one. The “large effect” label is a property of the point estimate, not of the evidence.
The typical-trial preset has the same problem more quietly: g = 0.6013 with an interval of −0.2253 to 1.4279.
The balanced 33/33 preset is the first of the four that excludes zero — 0.1293 to 1.1060 — and reallocating those same sixty-six people 60/6 puts zero back inside it. Sample size buys the interval; the point estimate is nearly free.
Why meta-analysis wants g rather than d
Hedges’ g is the standard input to a meta-analysis of continuous outcomes, and the reason is specific to pooling.
The bias in d does not average out. It is an upward bias in every study, so pooling twenty biased estimates produces a precisely estimated overstatement.
And it is worst in the smallest studies, which are also the ones most likely to be affected by publication bias — the two distortions push the same way.
The variance of g is what the inverse-variance weights are built from, so it has to be the corrected variance J²·Var(d), not Var(d) itself.
A pooled analysis of uncorrected d values with corrected weights, or the reverse, is a common and silent error. Both numbers are printed here so the pair being carried forward is unambiguous.
When g is the wrong statistic
g corrects one specific problem. Four other problems look similar and need something else.
If the two groups have very different spreads, use Glass’s delta, which divides by the control group’s SD alone rather than pooling two SDs that do not belong together.
If the data is skewed or has outliers, use Cliff’s delta. A single extreme value can flip the sign of d; it cannot move a rank-based measure more than 2/n.
If the outcome is ordinal, no standardised mean difference applies at all. The distance between “agree” and “strongly agree” is not a number.
If the measurements are paired, the pooled SD is the wrong denominator, and the correction factor uses a different degrees of freedom.
And if the groups differ on something other than the treatment, g corrects a sampling bias and leaves the confounding exactly where it was.
Reporting Hedges’ g
Four items, and the second is almost always missing.
Give g with its confidence interval. A bare effect size compared to a threshold is not a finding.
Give both group sizes, not the total. The allocation changes the interval by 71% at the same n, and a reader cannot reconstruct it from a total.
Say which correction you used. Exact and approximate J differ meaningfully below ten degrees of freedom, which is exactly where g gets used.
And give the raw means and SDs. A standardised effect is uninterpretable without them, and it lets a reader recompute anything on this page.
Sources and methodology
References for the bias correction and its variance.
Method. The correction factor is computed from its definition as a ratio of gamma functions, evaluated through log-gamma so it stays exact at large degrees of freedom, and the familiar 1 − 3/(4df − 1) shortcut is computed alongside it and printed rather than substituted. Cohen’s d is taken from the same shared routine the site’s effect size calculator uses, and the suite asserts agreement to 1e-14 on 200 generated designs, so the two pages cannot drift apart. The suite also asserts that J lies strictly between 0 and 1 for every admissible design, that g is never larger in magnitude than d, that J at df = 4 equals √(2/π) and at df = 3 equals √π/(2√1.5) in closed form, and that the shortcut’s error falls by a factor between 3.5 and 4.6 across five doublings of df — the signature of an O(1/df²) error. A group of one, or a pooled standard deviation of zero, returns no result. That engine is verified on every change against 134 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
Effect SizeCohen d, Hedges g and the overlap between groups, with a sample-size control that moves the p-value while leaving the effect size fixed — the same d gives t = 1.29 at n=30 and 23.57 at n=10,000.
Meta-AnalysisPool study effects by inverse variance under both fixed-effect and random-effects models, with τ², a prediction interval, per-study weights and leave-one-out influence.
Cliff's DeltaNon-parametric effect size from the full pairwise win–loss–tie count, with a DeLong interval, magnitude bands and Cohen's d for comparison.
Confidence IntervalIntervals for a mean or a proportion using t at every sample size and Wilson rather than the textbook Wald formula — with both methods shown, because Wald returns [0,0] at zero successes.
Sample SizeResponses needed for a target margin of error, with the finite-population correction and a table of the whole cost curve — because n scales with 1/margin², so the last point of precision costs more than the first ten.
t-testOne-sample, two-sample and paired t-tests defaulting to Welch, with Student's pooled version printed beside it — and a warning when the two disagree on the verdict.
An educational tool. Hedges’ g corrects a small-sample bias in the pooled standard deviation and nothing else — it cannot correct confounding, selection or measurement error. It also assumes the two groups are independent and that a pooled standard deviation is meaningful, which fails when the groups have very different spreads.