Math calculator

Games-Howell Test Calculator

All pairs compared, without assuming the spreads match.

All pairs, each with its own error term

The third group has five observations and about eight times the standard deviation of the other two — a variance ratio of 62.396. Tukey would use 42 pooled degrees of freedom for all three comparisons; the Welch degrees of freedom here are 37.93, 4.03 and 4.04.

3 groups, 3 comparisons, largest variance ratio 62.396×

0 of 3 pairs differ at α = 0.05

The largest variance is 62.40 times the smallest, which is well past where a pooled error term stops being defensible. Tukey HSD would have used 42 pooled degrees of freedom for every comparison; the Welch degrees of freedom here range from 4.03 to 37.93.

Variance ratio

62.396×

largest ÷ smallest

Comparisons

3

k(k−1)/2 for k = 3

Significant

0

at α = 0.05, family-wise

Tukey would use df

42

pooled, for every pair

Every pairwise comparison with its own standard error, degrees of freedom and interval
PairDifferenceSEqWelch dfp95% CI
Group 1 vs Group 2+0.05000.23990.208437.9320.988104[-0.778, 0.878]
Group 1 vs Group 3-1.95002.62820.74194.0320.863989[-15.148, 11.248]
Group 2 vs Group 3-2.00002.62870.76084.0350.857741[-15.196, 11.196]
Each group’s size, mean and variance
GroupnMeanVarianceSD
Group 12010.45001.10261.0501
Group 22010.40001.20001.0954
Group 3512.400068.80008.2946

What a pooled error term costs, measured

1,200 replications per row, three groups, ALL MEANS IDENTICAL — so every rejection below is a false positive and the correct rate is 5%.

Family-wise false positive rate for Tukey HSD and Games-Howell under three variance patterns
Group spreads and sizesTukey HSDGames-HowellTukey’s error
sd 1/1/1, n 15/15/154.67%4.75%none — both correct
sd 1/1/4, n 20/20/629.92%4.58%six times too many findings
sd 4/1/1, n 6/20/203.00%5.00%too few — real effects missed

Tukey fails in BOTH directions, and which one depends on whether the large variance sits in the small group or the large one — something you cannot know in advance of collecting the data. Games-Howell holds between 4.58% and 5.00% across all three, including the row where Tukey is correct.

Own SE and own df for every pair Studentized range, so the family rate is held Slightly liberal below about n = 6 per group

What this tool shows

With every group mean identical — so every rejection is a false positive — Tukey HSD fires 29.92% of the time when the large variance sits in the small group, and only 3.00% when it sits in the large one. Games-Howell holds between 4.58% and 5.00% across both, and across the case where variances are equal and Tukey is correct. Tukey fails in both directions, and which direction depends on something you cannot know before collecting the data.

  • Every pairwise comparison with its own standard error, not a pooled one
  • Welch-Satterthwaite degrees of freedom per pair, printed beside Tukey’s pooled value
  • Studentized range p-values and simultaneous confidence intervals
  • The largest variance ratio in your data, as the flag for whether pooling was safe
  • A measured table of Tukey’s false-positive rate under three variance patterns
  • Group sizes, means and variances, so the source of any df collapse is visible
All pairs Welch df per pair No pooling Measured calibration

Slightly liberal below about six observations per group.

Updated 12 September 2026 · Works in any browser, no installation

The Games-Howell test compares every pair of groups while giving each comparison its own standard error and its own degrees of freedom, so unequal variances and unequal group sizes do not corrupt it. It is the post-hoc counterpart of a Welch t-test, in the same way that Tukey HSD is the counterpart of a pooled-variance t-test — and it holds the family-wise error rate using the same studentized range distribution.

At a glance

Formula shown
For each pair, SE = √[(s²ᵢ/nᵢ + s²ⱼ/nⱼ)/2] and q = |x̄ᵢ − x̄ⱼ|/SE, compared against the studentized range distribution on k groups and Welch-Satterthwaite degrees of freedom df = (s²ᵢ/nᵢ + s²ⱼ/nⱼ)² / [(s²ᵢ/nᵢ)²/(nᵢ−1) + (s²ⱼ/nⱼ)²/(nⱼ−1)]. Tukey HSD uses one pooled MS_error and one df for every pair; here both change from pair to pair, which is the whole difference.
Scenario support
Post-hoc comparisons after any ANOVA where a Levene test flags unequal variances, groups of very different sizes, outcomes whose spread grows with their mean, comparisons involving one small pilot group against larger ones, and any design where a treatment plausibly changes variability as well as level.
Educational estimate
Planning support from the values you enter — not professional advice.

Tukey fails in both directions, and you cannot predict which

The usual warning is that Tukey is “liberal” with unequal variances. That is half true, and the half it leaves out is the half that costs findings.

Three groups, all with the SAME mean, 1,200 replications per pattern. Every rejection below is a false positive and the correct rate is 5%.

Equal spreads and equal n: Tukey 4.67%, Games-Howell 4.75%. Both correct. Using Games-Howell where Tukey applies costs essentially nothing.

Large variance in the SMALL group (sd 1/1/4, n 20/20/6): Tukey 29.92%, Games-Howell 4.58%. Six times too many findings. The pooled error term is dominated by the two large, precise groups, so the small noisy group is compared against a standard error far smaller than its own.

Large variance in the LARGE group (sd 4/1/1, n 6/20/20): Tukey 3.00%, Games-Howell 5.00%. Now the pooled term is inflated by the big noisy group, and Tukey becomes conservative — missing real differences between the two precise groups.

Which direction you get depends on where the variance sits relative to the sample sizes, and that is not something a study design can guarantee in advance. Games-Howell holds between 4.58% and 5.00% in all three.

Fractional degrees of freedom are the mechanism, not a rounding artefact

The Welch-Satterthwaite degrees of freedom are almost never a whole number, and their collapse is what protects the error rate.

On the tool’s first preset, Tukey would use 42 pooled degrees of freedom for every comparison. Games-Howell uses 37.93 for the pair of large precise groups and 4.03 and 4.04 for the pairs involving the small noisy one.

Four degrees of freedom against forty-two is a completely different reference distribution. That is the difference between a critical value that protects the rate and one that does not.

The df depend on both the variances and the sizes, so a group can lose degrees of freedom for being noisy, for being small, or for both — and the tool flags any pair whose df fall below half the pooled value.

The studentized range distribution is perfectly well defined at fractional df, so nothing is being approximated at that step. The approximation is in Welch’s formula itself, which is what makes the procedure very slightly liberal at tiny sample sizes.

Games-Howell, Tukey, Tamhane and Dunnett

Four procedures that all control a family-wise rate, differing in what they assume and which comparisons they make.

Tukey HSD is the most powerful when its assumptions hold — equal variances, ideally equal n — and is the right default when a Levene test gives no reason to doubt them.

Tamhane’s T2 is the conservative alternative, built on a Sidak bound rather than the studentized range. It also handles unequal variances and gives up more power doing so, which makes it the safer choice only when groups are very small.

Dunnett’s test answers a different question — every treatment against one control rather than all pairs — and is far more powerful for that question because it runs k comparisons instead of k(k−1)/2.

A Bonferroni correction on Welch t-tests is the simplest unequal-variance route and is more conservative than Games-Howell, because it treats the comparisons as independent when they share groups.

The practical rule: run Levene first. Equal variances, use Tukey; unequal, use Games-Howell; comparing against a control, use Dunnett regardless. And since Games-Howell costs essentially nothing when variances are equal, defaulting to it is defensible.

Where Games-Howell itself breaks down

It is not a universal solvent. Its own approximation degrades in a specific and predictable place.

Below about six observations per group it is slightly liberal, producing a family-wise rate a little above the nominal level. The Welch degrees of freedom are an approximation, and approximations of a df are worst when the df are small.

The variance estimates themselves are unstable at small n. A sample variance from five observations is a noisy estimate, and the whole procedure is built on per-group variances.

It still assumes normality within groups. Unequal variances are handled; heavy tails and skew are not, and the rank-based Kruskal-Wallis test with its own post-hoc is the route for those.

And it still assumes independent observations, which no post-hoc procedure can repair — repeated measurements on the same subjects need a within-subjects design instead.

It does not need a significant ANOVA first

The traditional workflow runs an omnibus F test and only proceeds to post-hoc comparisons if it is significant. That gate is inappropriate here for two separate reasons.

The F test assumes equal variances too. Gating an unequal-variance procedure behind a test that assumes equal variances puts the assumption back in through the side door.

Games-Howell controls its own family-wise rate. It does not rely on the omnibus test for protection, so the gate adds a condition without adding a guarantee.

The gate also costs power. A non-significant omnibus F with one real pairwise difference is entirely possible, especially with more than three groups, and the gate discards it.

Where an omnibus test is genuinely wanted with unequal variances, use Welch’s ANOVA, which is the matching omnibus procedure — but Games-Howell does not require it to have been run.

Reporting a Games-Howell test

Four things, and the second is what tells a reader why this procedure rather than Tukey.

Name the procedure. “Post-hoc comparisons” is not a method; Tukey, Games-Howell and Bonferroni give different answers on the same data.

Give the group variances or standard deviations. They are the justification for not pooling, and they take one row of a table.

Report the fractional degrees of freedom per comparison. They are not a typo, they differ between pairs, and they are the mechanism by which the procedure works.

Report every pair, including the ones that were not significant. A family-wise correction applied to a family and then reported selectively is not a correction at all.

Sources and methodology

References for Games-Howell and the unequal-variance problem.

Method. Each pair gets its own standard error from the two group variances and its own Welch-Satterthwaite degrees of freedom, and the critical studentized range is found per pair by bisection on the studentized range distribution at those fractional degrees of freedom — memoised, since a calibration sweep over thousands of datasets would otherwise take minutes rather than seconds. The calibration table on this page is a measurement rather than a citation: 1,200 replications per variance pattern with every group drawn from a distribution with the SAME mean, so every rejection counted is a false positive and the target is the nominal 5%. Tukey HSD is run on the identical datasets through the site’s own Tukey engine, which is what makes the 29.92% against 4.58% and the 3.00% against 5.00% a like-for-like comparison. Fewer than two groups, a group with one value and a group with no variation all return no result. That engine is verified on every change against 96 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

Tukey HSDEvery pairwise comparison after an ANOVA with simultaneous confidence intervals — and the studentized range integrated rather than interpolated from a table, so any group count, df and level works. At two groups the critical q is exactly √2 times the critical t.
Levene's TestComputes both centre conventions side by side and shows the median is what makes it robust: on skewed data the mean-centred original rejects 15.00% of true nulls against the median version's 4.47%.
Dunnett's TestMany treatments against one control, with the critical value computed by numerical integration rather than read from a table — at one comparison it collapses to the two-sided t to 3.55e-15.
One-Way ANOVAThe full F table with eta and omega squared, plus every pairwise gap — because a significant F says something differs and never says which, and ten groups tested pairwise carry a 90% false-positive rate.
Bonferroni CorrectionAdjusts p-values by Bonferroni, Holm, Šidák, Holm-Šidák and Benjamini-Hochberg at once — and shows that Holm controls exactly what Bonferroni controls while never rejecting fewer, which makes plain Bonferroni dominated.
Kruskal-WallisApplies the tie correction and shows it against the uncorrected value, because on ordinal data it moves p from 0.054 to 0.027 — across the conventional threshold, on identical data.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool. Games-Howell handles unequal variances and unequal group sizes but still assumes independent, roughly normal observations within each group, and it is slightly liberal below about six observations per group because the Welch degrees of freedom are themselves an approximation.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (5 updates)

Published 12 September 2026

  1. Published an all-pairs post-hoc test that gives every comparison its own standard error and its own Welch-Satterthwaite degrees of freedom instead of pooling.
  2. Measured what pooling costs, and found the usual warning is only half right. Across 1,200 replications per pattern with every group mean IDENTICAL — so every rejection is a false positive against a nominal 5% — Tukey HSD fires 29.92% of the time when the large variance sits in the small group and only 3.00% when it sits in the large one. Games-Howell holds between 4.58% and 5.00% in both, and in the equal-variance control where Tukey is correct.
  3. Which means Tukey fails in BOTH directions, and which direction depends on where the variance sits relative to the sample sizes — something a study design cannot guarantee in advance.
  4. Showed the mechanism on the page: Tukey would use 42 pooled degrees of freedom for every comparison on the shipped preset, while the Welch values are 37.93, 4.03 and 4.04.
  5. Verified that a pair is significant exactly when its confidence interval excludes zero, across 100 generated datasets, and that Welch degrees of freedom never exceed the pooled degrees of freedom they replace.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.