Three hundred and twenty predicted probabilities with their outcomes. At the conventional ten groups the statistic is 10.5360 on 8 degrees of freedom, p = 0.2294 — comfortably calibrated. At nine groups it is p = 0.0284, which most readers would call a failure. At six groups it is 0.9583. Same forecasts, same outcomes, three answers spanning almost the whole unit interval. Hosmer and Lemeshow chose ten groups by convention rather than from theory, and the statistic has no asymptotic justification for any particular choice. Reporting one p-value from one group count hides this entirely.
320 cases · 10 risk groups · 8 degrees of freedom
χ² 10.5360, p = 0.2294
Across group counts from five to fifteen the p-value runs from 0.0284 to 0.9583, crossing 0.05 on the way. The group count is a convention rather than a property of the data, so a single p-value from a single choice is not a reportable result here — the sweep below is.
p at this group count
0.2294
10 groups, 8 df
Across the sweep
0.0284 to 0.9583
five to fifteen groups
Verdict changes
yes
the binning decides it
Worst gap
0.1232
group 2, promised against observed
The p-value against the group count
The statistic and p-value at every group count from five to fifteen
Groups
χ²
df
p
Verdict at 0.05
5
1.8091
3
0.6130
no evidence against
6
0.6420
4
0.9583
no evidence against
7
7.9526
5
0.1589
no evidence against
8
9.3602
6
0.1543
no evidence against
9
15.6658
7
0.0284
miscalibrated
10
10.5360
8
0.2294
no evidence against
11
12.9920
9
0.1630
no evidence against
12
12.2137
10
0.2710
no evidence against
13
15.9348
11
0.1436
no evidence against
14
19.9249
12
0.0685
no evidence against
15
22.4550
13
0.0487
miscalibrated
None of these group counts is more correct than any other. Ten is convention, not theory. A result that holds across the whole column is worth reporting; one that depends on which row you read is a property of the binning rather than of the model.
Group by group
Each risk group with the events it was promised and the events it had
Group
Cases
Risk range
Promised
Observed
Gap
χ² share
1
32
0.001 to 0.096
0.0455
0.0938
0.0482
16.3%
2
32
0.104 to 0.202
0.1545
0.0313
-0.1232
35.3%
3
32
0.202 to 0.303
0.2481
0.2813
0.0331
1.8%
4
32
0.308 to 0.409
0.3550
0.3750
0.0200
0.5%
5
32
0.414 to 0.542
0.4911
0.5938
0.1027
12.8%
6
32
0.543 to 0.621
0.5798
0.5625
-0.0173
0.4%
7
32
0.624 to 0.706
0.6723
0.6563
-0.0160
0.4%
8
32
0.709 to 0.811
0.7611
0.7813
0.0202
0.7%
9
32
0.813 to 0.901
0.8592
0.7813
-0.0779
15.2%
10
32
0.902 to 0.995
0.9481
1.0000
0.0519
16.6%
The gap column is what calibration means: a model promising 30% risk to a group that had 45% events is wrong by 15 points there, whatever the p-value says. Reading these directly is more informative than the test, and it is the part that survives a change in the group count.
The Brier decomposition on the same forecasts gives a reliability of 0.00348 against a resolution of 0.08969. Reliability is calibration measured on a continuous scale rather than tested; resolution is discrimination, which this test says nothing about at all.
The test needs roughly ten expected events in every group, which at ten groups means several hundred cases. On a small sample a badly miscalibrated model returns a comfortable p-value, which is what the fourth preset demonstrates rather than describes.
Group count swept Per-group gaps Brier alongside Calibration only, not skill
What this tool shows
On the shipped preset the same three hundred and twenty forecasts give p = 0.9583 at six groups and p = 0.0284 at nine. The conventional ten gives 0.2294. Hosmer and Lemeshow chose ten by convention, the statistic has no asymptotic justification for any particular number, and the p-value moves with that arbitrary choice — here across 0.05. So this page reports the whole sweep from five groups to fifteen rather than one number from one of them.
The Hosmer-Lemeshow statistic at your chosen group count, with its degrees of freedom
The same statistic at every group count from five to fifteen, and whether the verdict changes
Per-group promised and observed rates, with each group’s share of the total statistic
The Brier reliability on the same forecasts — calibration measured rather than tested
A preset where a badly miscalibrated model on forty cases returns a comfortable p-value
Why a non-significant result is not evidence of calibration
Group count swept Per-group gaps Brier alongside Verdict stability
It tests calibration. It says nothing about discrimination.
Updated 13 September 2026 · Works in any browser, no installation
The Hosmer-Lemeshow test sorts cases by predicted risk, splits them into equal-sized groups, and compares how many events each group was promised against how many it had. It answers one narrow question — are the predicted probabilities right on average within each risk band — and answers it with a p-value that depends on how many bands you chose, which is a choice nothing in the data determines.
At a glance
Formula shown
Sort by predicted probability, split into g groups of roughly equal size, and sum (Oₖ − Eₖ)²/Eₖ over both outcome cells of every group, where Eₖ is the sum of predicted probabilities in group k. Refer the total to chi-square on g − 2 degrees of freedom. The two subtracted degrees of freedom are conventional rather than derived, and the whole reference distribution is an approximation justified by simulation rather than by asymptotic theory.
Scenario support
Checking whether a logistic model’s predicted probabilities match observed rates, auditing a risk score before it is used for decisions, comparing calibration between a development and a validation sample, and any setting where a predicted probability will be read as a probability.
Educational estimate
Planning support from the values you enter — not professional advice.
The group count decides the answer
This is the reason the page exists in this form, and it is easy to demonstrate and rarely mentioned.
At six groups: p = 0.9583. At nine: p = 0.0284. Same forecasts, same outcomes, nothing changed but the binning.
The conventional ten gives 0.2294, which is the number almost every write-up would report, with no indication that the neighbouring choices disagree.
Ten is a convention from the original 1980 paper, chosen because it worked in simulations, not because the asymptotics require it.
Which makes a single p-value here uninterpretable. A verdict that survives the whole sweep — as on the second and third presets — means something; one that depends on the row you read is a property of the binning.
It needs a large sample, and fails quietly without one
The fourth preset is a model that is obviously wrong and a test that does not notice, which is the more dangerous of the two failure modes.
Forty cases, forecasts averaging 0.4580, observed rate 0.3250. The model promises forty-six per cent risk and delivers thirty-three.
The test returns p = 0.3469. Comfortably calibrated, by the usual reading.
The largest gap between promised and observed rate in any group is 0.5573, which is enormous and still not enough.
The requirement is roughly ten expected events per group, which at ten groups and a 30% event rate means several hundred cases. Below that, a non-significant result carries almost no information.
Read the gaps, not the p-value
The per-group table is the part of the output that survives every objection to the test, and it is usually discarded in favour of the single number.
A group promised 30% risk that had 45% events is wrong by fifteen points there. That is a fact about the model, independent of any reference distribution.
The direction matters as much as the size. Under-prediction at high risk and over-prediction at low risk are different failures with different consequences.
So does where it happens. A model miscalibrated only in its top risk group may be unusable for exactly the decisions it was built for, while the overall statistic barely moves.
The χ² share column names the culprit. When one group contributes most of the statistic, the problem is local and the remedy is local — a missing interaction, a non-linear term — rather than a wholesale recalibration.
Calibration is not discrimination
A model can pass this test comfortably and be worthless, which is the most important thing the test does not tell you.
Predict the base rate for everybody and calibration is perfect. Every group is promised the same rate and every group delivers roughly that rate.
Such a model has no discrimination at all. It separates nobody from anybody, and the Hosmer-Lemeshow test is content.
The Brier decomposition separates the two, and this page prints both terms: reliability is calibration, resolution is discrimination, and a flat forecast scores zero on the second.
The other half of the picture is the AUC, which measures ranking and is completely insensitive to calibration — a model can have an AUC of 0.9 with every probability twice what it should be.
Why the test is out of favour
Hosmer and Lemeshow themselves have written about its limitations, and the case against it is worth stating plainly on a page that computes it.
The grouping is arbitrary, which is the point demonstrated above.
The reference distribution is approximate. The g − 2 degrees of freedom come from simulation rather than from theory, and the approximation degrades when groups are small.
At large samples it rejects almost everything. Any model is slightly miscalibrated, and with ten thousand cases the test finds it — which makes a significant result uninformative at exactly the sample sizes where the test is otherwise valid.
The modern alternatives are graphical and continuous. A calibration curve with a smoother, the calibration slope and intercept, or the Brier reliability term — none of which need a group count and all of which report a magnitude rather than a verdict.
When it is still the right tool
Despite all of that, there are settings where this test earns its place, and dismissing it entirely goes too far.
It is expected in clinical prediction literature, so reporting it alongside better measures is usually the pragmatic choice rather than a concession.
At a few hundred cases it has genuine power without the large-sample hypersensitivity — which is a narrow but real window.
The per-group table is a useful diagnostic in its own right, independent of the p-value attached to it.
And the sweep on this page turns a weakness into information. A verdict stable across every group count from five to fifteen is worth more than any single p-value, and an unstable one tells you the evidence is thin in a way one number would conceal.
Reporting a calibration check
Four items, and the first is what the whole page is about.
Give the group count. A p-value without it is not reproducible, and the neighbouring choices may disagree.
Give the sample size and the number of events. Below a few hundred events, a comfortable p-value means the test had no power rather than that the model is calibrated.
Give the largest promised-against-observed gap. It is a magnitude rather than a verdict, and it does not move when the binning does.
And report a discrimination measure too. Calibration alone cannot distinguish a useful model from one that predicts the base rate for everybody.
Method. Cases are sorted by predicted probability and split into groups of as near equal size as the sample allows, which is the deciles-of-risk form of the test rather than the fixed-cutpoint variant. A group whose expected count in one outcome cell is effectively zero contributes nothing rather than an infinity, and the degrees of freedom are reduced to match, so a degenerate group cannot silently produce a meaningless statistic. The sweep from five groups to fifteen runs on every input because the group count is a convention rather than a property of the data, and the page reports whether the verdict at 0.05 changes anywhere within it. The verification suite asserts the structural facts on sixty generated sets — degrees of freedom equal to groups minus two, a non-negative statistic with a valid p-value, and a sweep of exactly eleven entries each with matching degrees of freedom — and separately checks that the group counts sum back to the sample size. That engine is verified on every change against 100 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
Brier ScoreThe Brier score split into reliability, resolution and uncertainty, with the term the decomposition drops printed rather than absorbed, plus a skill score and a calibration table.
Logistic RegressionLogistic regression with odds ratios converted to risk ratios at your own event rate, a likelihood ratio test, AUC, and separation reported rather than hidden.
ROC Curve & AUCBuilds the curve from raw scores with every threshold enumerated, and computes the AUC twice — trapezoid and Mann-Whitney U — which agree to 1.11e-16 across 300 datasets.
Log LossLog loss with the Brier score on the same forecasts, the penalty each applies at increasing confidence, and how much of the total a single case carries.
Chi-SquareGoodness of fit and tests of independence with every expected count and per-cell contribution shown — because the validity condition is about expected counts, not observed ones, and most calculators hide them.
Odds Ratio to Risk RatioConvert an odds ratio to a risk ratio at any baseline risk, with the interval converted too and the same odds ratio shown across eleven baselines.
An educational tool. The Hosmer-Lemeshow p-value depends on the number of groups, which is a convention rather than a property of the data — on the shipped preset it ranges from 0.03 to 0.96 across ordinary choices. The test needs roughly ten expected events per group, so a comfortable result on a small sample means it had no power; and it tests calibration only, which a model predicting the base rate for everyone passes perfectly.