Math calculator

Bland-Altman Calculator

Do two methods agree?

Limits of agreement

Twenty-five subjects measured twice. The correlation is 0.95550, which reads as excellent agreement and is not agreement at all: the limits run from −46.92 to +42.22, a span of 89.14 — 19.1% of the mean measurement. On any one subject the two instruments can differ by that much. The bias itself is −2.35 with an interval of −11.73 to +7.04, so there is no systematic offset to correct.

25 subjects · bias -2.3480

Limits of agreement -46.9165 to 42.2205

On any one subject the two methods can differ by that much. Pearson's r on the same data is 0.95550 — which would read as excellent agreement and is not agreement at all: a correlation asks whether the methods move together, not whether they match. The bias of -2.3480 has its own interval of -11.7342 to 7.0382, which includes zero.

Bias

-2.3480

CI -11.734 to 7.038

Limits of agreement

-46.916 to 42.220

span 89.137

Pearson's r

0.95550

not a measure of agreement

Proportional bias

0.0224

p = 0.7213

The limits are estimates too. The upper limit is 42.2205 and its own 95% interval runs 25.9608 to 58.4801 — a span of 32.5193 at 25 subjects. Almost no published plot draws them, and they are what tells you whether the study was large enough.

The plot

-500300400500600average of the two methodsdifference (A − B)

1 of 25 points fall outside the limits, against roughly 1.3 expected. A line that slopes means the bias depends on the measurement level, and a single pair of limits then describes no part of the range.

Every subject

Each subject’s two measurements, their average and their difference
#Method AMethod BAverageDifference
1433.2000492.4000462.8000-59.2000 — outside
2331.4000322.8000327.1000+8.6000
3498.8000487.0000492.9000+11.8000
4553.4000553.7000553.5500-0.3000
5409.3000388.8000399.0500+20.5000
6408.1000445.1000426.6000-37.0000
7596.1000575.3000585.7000+20.8000
8393.8000425.7000409.7500-31.9000
9479.3000498.4000488.8500-19.1000
10498.3000533.7000516.0000-35.4000
11451.3000444.3000447.8000+7.0000
12672.4000664.0000668.2000+8.4000
13426.1000393.4000409.7500+32.7000
14416.0000430.2000423.1000-14.2000
15516.4000484.2000500.3000+32.2000
16408.4000415.8000412.1000-7.4000
17493.6000508.8000501.2000-15.2000
18413.4000419.0000416.2000-5.6000
19506.4000499.9000503.1500+6.5000
20375.5000367.8000371.6500+7.7000
21421.2000407.2000414.2000+14.0000
22512.7000533.1000522.9000-20.4000
23511.3000500.8000506.0500+10.5000
24386.7000397.4000392.0500-10.7000
25538.7000521.7000530.2000+17.0000
Intervals on the limits Proportional-bias test r shown, to be argued against Limits are not a pass mark

What this tool shows

The shipped study correlates at r = 0.95550 and its limits of agreement span 89.14 — 19.1% of the mean measurement. On any one subject the two methods can differ by that much. A correlation asks whether two methods move together; agreement asks whether they match, and the two questions come apart completely. The limits have their own confidence intervals here too, which almost no published plot draws.

  • Bias and limits of agreement from paired measurements
  • Confidence intervals on the limits themselves, not only on the bias
  • A regression test for proportional bias — whether the difference grows with the measurement
  • Pearson’s r printed alongside, because it is the number this analysis exists to displace
  • The plot against the average of the two methods, never against either one
  • Every subject listed with its average and difference, and which points fall outside
Intervals on the limits Proportional-bias test r shown for contrast Per-subject table

The limits describe the data. Whether they are acceptable is a clinical judgement.

Updated 13 September 2026 · Works in any browser, no installation

Bland-Altman plots the difference between two methods against their average, and draws limits two standard deviations either side of the mean difference. Those limits are the answer: they say how far apart the two methods can be on a single subject, which is the question anyone deciding whether to swap one for the other actually has. A correlation cannot answer it — two methods that differ by a factor of ten correlate perfectly.

At a glance

Formula shown
For each subject, difference dᵢ = Aᵢ − Bᵢ and average mᵢ = (Aᵢ + Bᵢ)/2. Bias = d̄, and the limits are d̄ ± 1.96·s_d. The bias has standard error s_d/√n; each limit has the larger standard error s_d·√(1/n + 1.96²/(2(n−1))), which is why the limits are far less certain than the bias and why an interval on them is worth printing. Proportional bias is tested by regressing dᵢ on mᵢ: a non-zero slope means a single pair of limits describes no part of the range.
Scenario support
Validating a new instrument against a reference, comparing a point-of-care device with a laboratory assay, checking whether two observers can be treated as interchangeable, assessing test-retest repeatability of a single method, and any decision about replacing one measurement process with another.
Educational estimate
Planning support from the values you enter — not professional advice.

r = 0.956, and the methods differ by 89 units

Bland and Altman’s 1986 paper existed to stop people using a correlation for this, and the shipped preset shows why in one line.

Pearson’s r on those twenty-five subjects is 0.95550. Reported on its own, that reads as excellent agreement.

The limits of agreement run −46.92 to +42.22, a span of 89.14. The mean measurement is around 466, so on any single subject the two instruments can disagree by 19.1% of the quantity being measured.

A correlation measures whether the methods move together, not whether they match. Multiply one method by ten and the correlation is unchanged; the agreement is destroyed.

It also rises with the range of the sample. Measure a wider spread of subjects and r goes up with no change in the instruments at all, which makes it trivially inflatable by choosing who to recruit.

The limits are estimates, and nobody draws their intervals

The limits are printed as though they were known quantities. They are two numbers estimated from the same n subjects as everything else.

On the shipped preset the upper limit is 42.22. Its own 95% confidence interval runs from 25.96 to 58.48.

That is a span of 32.52 on a limit of 42.22 — the limit is uncertain by about three quarters of its own value, at twenty-five subjects.

The standard error of a limit is much larger than that of the bias because it carries the uncertainty in the standard deviation as well as in the mean. That is why a study can pin down the bias and still say almost nothing about the limits.

Bland and Altman recommended at least fifty subjects for this reason, and most published studies use fewer.

When one pair of limits describes nothing

The whole method assumes the difference between the two instruments does not depend on how large the measurement is. When it does, the limits are an average of two different situations.

The proportional-bias preset regresses difference on average and gets a slope of −0.2168 with p below 0.00001. The second method reads about 25% high.

So at the bottom of the range the two nearly agree and at the top they are 30 units apart. A single pair of limits sits between those and describes neither.

The usual fix is to work on the logarithms, which turns a constant ratio into a constant difference and makes the limits interpretable as percentages.

The alternative is to fit the relationship rather than average it away — a Deming or Passing-Bablok regression estimates the slope and intercept between the two methods directly.

Why the x axis is the average and not the reference

It is tempting to plot the difference against the reference method, on the grounds that the reference is the truth. Doing so manufactures a slope out of nothing.

The difference A − B contains B, and so does the x axis. Any measurement error in B appears on both axes with opposite signs, which produces a downward slope even when the two methods agree perfectly on average.

Using the average splits that error between the axes and removes most of the artefact.

It is not a perfect fix — the average still contains both errors — but it is unbiased when the two methods have similar error variances, which is the usual case in a method comparison.

If one method genuinely is a reference standard with negligible error, plotting against it is defensible; the assumption then needs stating, because it is rarely true.

The limits do not say whether the methods are interchangeable

This is the step the statistics cannot take, and the one most often skipped.

Limits of ±44 are fine for some measurements and catastrophic for others. The analysis produces a number; whether it is acceptable is a question about consequences.

The acceptable difference has to be set before the study, from what a wrong reading would cause, not from what the data happened to produce.

Then the comparison is between that limit and the confidence interval on the limits — not against the point estimates, which is why those intervals are printed here.

A study that cannot exclude an unacceptable limit has not shown agreement, even if its point estimates look comfortable.

Two methods cannot agree better than one method agrees with itself

A wide set of limits is often read as a fault in the new method. It may be a fault in the reference.

Every method has its own repeatability. Measure the same subject twice with one instrument and the two readings differ.

The variance of the difference between two methods includes both. So the limits can never be narrower than the repeatability of the worse of the two.

Which means a repeatability study should come first. Run each method twice on the same subjects and put those differences through this same analysis.

If the limits between methods are close to the repeatability limits within a method, the two instruments agree as well as the measurement process allows, and no better instrument would help.

Reporting a Bland-Altman analysis

Five items, and the second and third are the ones that go missing.

Give the bias and the limits, with n. All three, and to enough figures that a reader can recompute the standard deviation.

Give confidence intervals on the limits. Not only on the bias — the limits are the result, and they are the less certain half.

Report the proportional-bias test. A sloping plot invalidates the limits, and eyeballing the scatter is not a test.

State the acceptable difference you set in advance, and say whether the limits clear it.

And do not report a correlation as evidence of agreement. If you report one, say what it is for.

Sources and methodology

References for limits of agreement and their precision.

Method. Differences are taken as A minus B and plotted against the average of the two, never against either method alone — plotting against a supposed reference puts that method’s measurement error on both axes with opposite signs and manufactures a slope out of nothing. The limits use a 1.96 multiplier on the standard deviation of the differences, and each limit carries its own confidence interval built on the larger standard error that accounts for uncertainty in the standard deviation as well as in the mean. Proportional bias is tested by an ordinary regression of difference on average, which is reported with its p-value rather than left to the eye. Pearson’s r is computed on the same pairs and printed, because it is the statistic this analysis exists to displace and the contrast is the point. The suite asserts that the bias is always exactly the difference of the two means, that the limits sit exactly 1.96 standard deviations either side of it, and that swapping the two methods flips the bias and leaves the spread untouched — all on 200 generated studies. That engine is verified on every change against 103 assertions. The count and the per-case breakdown are published on the formula verification page.

Related calculators

Where this goes next:

Concordance CorrelationLin's concordance correlation with its exact decomposition into precision and accuracy, the scale and location shifts separated, and a bounded interval.
Passing-BablokRobust method-comparison regression with rank-based intervals, the symmetry shift reported, and the pairwise-slope distribution shown.
Deming RegressionRegression when both variables carry error, with a settable error-variance ratio, jackknife intervals and least squares shown as the limiting case.
Intraclass CorrelationAll six ICC forms from one subject-by-rater matrix, with the rater means that drive them apart: one 8x3 matrix gives ICC(1,1) = 0.1277 and ICC(3,1) = 0.9852.
Gauge R&RANOVA gauge repeatability and reproducibility with both acceptance criteria, the operator-by-part interaction tested, and the full variance decomposition.
Correlation CoefficientReports Pearson, Spearman and Kendall together with the scatter plot, and ships Anscombe's quartet built in — four datasets with an identical r of 0.816 that Spearman tells apart.

More in Math, or browse all calculators.

Educational use disclaimer

An educational tool. Limits of agreement describe the data and do not say whether two methods are interchangeable — that requires an acceptable difference set before the study from what a wrong reading would cause. The limits also assume the differences are roughly normal and that the bias does not depend on the measurement level; the proportional-bias test on this page checks the second, and a sloping result invalidates the limits.

How we calculate · Found an error? email us

Authorship & verification

Written and maintained by , a business operator who builds spreadsheet-based calculators.

What's changed (5 updates)

Published 13 September 2026

  1. Launched limits-of-agreement analysis with the difference plotted against the average of the two methods, never against either alone.
  2. Added confidence intervals on the limits themselves — the upper limit of 42.22 on the shipped study has its own interval from 25.96 to 58.48.
  3. Added a regression test for proportional bias, since a sloping difference plot invalidates a single pair of limits.
  4. Printed Pearson r beside the limits for contrast: 0.95550 on a study whose limits span 19.1% of the mean measurement.
  5. Listed every subject with its average and difference, and flagged which points fall outside the limits.

Add this calculator to your site

Responsive embed — and private: nothing your visitors type leaves their browser.