Twenty-five subjects measured twice. The correlation is 0.95550, which reads as excellent agreement and is not agreement at all: the limits run from −46.92 to +42.22, a span of 89.14 — 19.1% of the mean measurement. On any one subject the two instruments can differ by that much. The bias itself is −2.35 with an interval of −11.73 to +7.04, so there is no systematic offset to correct.
25 subjects · bias -2.3480
Limits of agreement -46.9165 to 42.2205
On any one subject the two methods can differ by that much. Pearson's r on the same data is 0.95550 — which would read as excellent agreement and is not agreement at all: a correlation asks whether the methods move together, not whether they match. The bias of -2.3480 has its own interval of -11.7342 to 7.0382, which includes zero.
Bias
-2.3480
CI -11.734 to 7.038
Limits of agreement
-46.916 to 42.220
span 89.137
Pearson's r
0.95550
not a measure of agreement
Proportional bias
0.0224
p = 0.7213
The limits are estimates too. The upper limit is 42.2205 and its own 95% interval runs 25.9608 to 58.4801 — a span of 32.5193 at 25 subjects. Almost no published plot draws them, and they are what tells you whether the study was large enough.
The plot
1 of 25 points fall outside the limits, against roughly 1.3 expected. A line that slopes means the bias depends on the measurement level, and a single pair of limits then describes no part of the range.
Every subject
Each subject’s two measurements, their average and their difference
#
Method A
Method B
Average
Difference
1
433.2000
492.4000
462.8000
-59.2000 — outside
2
331.4000
322.8000
327.1000
+8.6000
3
498.8000
487.0000
492.9000
+11.8000
4
553.4000
553.7000
553.5500
-0.3000
5
409.3000
388.8000
399.0500
+20.5000
6
408.1000
445.1000
426.6000
-37.0000
7
596.1000
575.3000
585.7000
+20.8000
8
393.8000
425.7000
409.7500
-31.9000
9
479.3000
498.4000
488.8500
-19.1000
10
498.3000
533.7000
516.0000
-35.4000
11
451.3000
444.3000
447.8000
+7.0000
12
672.4000
664.0000
668.2000
+8.4000
13
426.1000
393.4000
409.7500
+32.7000
14
416.0000
430.2000
423.1000
-14.2000
15
516.4000
484.2000
500.3000
+32.2000
16
408.4000
415.8000
412.1000
-7.4000
17
493.6000
508.8000
501.2000
-15.2000
18
413.4000
419.0000
416.2000
-5.6000
19
506.4000
499.9000
503.1500
+6.5000
20
375.5000
367.8000
371.6500
+7.7000
21
421.2000
407.2000
414.2000
+14.0000
22
512.7000
533.1000
522.9000
-20.4000
23
511.3000
500.8000
506.0500
+10.5000
24
386.7000
397.4000
392.0500
-10.7000
25
538.7000
521.7000
530.2000
+17.0000
Intervals on the limits Proportional-bias test r shown, to be argued against Limits are not a pass mark
What this tool shows
The shipped study correlates at r = 0.95550 and its limits of agreement span 89.14 — 19.1% of the mean measurement. On any one subject the two methods can differ by that much. A correlation asks whether two methods move together; agreement asks whether they match, and the two questions come apart completely. The limits have their own confidence intervals here too, which almost no published plot draws.
Bias and limits of agreement from paired measurements
Confidence intervals on the limits themselves, not only on the bias
A regression test for proportional bias — whether the difference grows with the measurement
Pearson’s r printed alongside, because it is the number this analysis exists to displace
The plot against the average of the two methods, never against either one
Every subject listed with its average and difference, and which points fall outside
Intervals on the limits Proportional-bias test r shown for contrast Per-subject table
The limits describe the data. Whether they are acceptable is a clinical judgement.
Updated 13 September 2026 · Works in any browser, no installation
Bland-Altman plots the difference between two methods against their average, and draws limits two standard deviations either side of the mean difference. Those limits are the answer: they say how far apart the two methods can be on a single subject, which is the question anyone deciding whether to swap one for the other actually has. A correlation cannot answer it — two methods that differ by a factor of ten correlate perfectly.
At a glance
Formula shown
For each subject, difference dᵢ = Aᵢ − Bᵢ and average mᵢ = (Aᵢ + Bᵢ)/2. Bias = d̄, and the limits are d̄ ± 1.96·s_d. The bias has standard error s_d/√n; each limit has the larger standard error s_d·√(1/n + 1.96²/(2(n−1))), which is why the limits are far less certain than the bias and why an interval on them is worth printing. Proportional bias is tested by regressing dᵢ on mᵢ: a non-zero slope means a single pair of limits describes no part of the range.
Scenario support
Validating a new instrument against a reference, comparing a point-of-care device with a laboratory assay, checking whether two observers can be treated as interchangeable, assessing test-retest repeatability of a single method, and any decision about replacing one measurement process with another.
Educational estimate
Planning support from the values you enter — not professional advice.
r = 0.956, and the methods differ by 89 units
Bland and Altman’s 1986 paper existed to stop people using a correlation for this, and the shipped preset shows why in one line.
Pearson’s r on those twenty-five subjects is 0.95550. Reported on its own, that reads as excellent agreement.
The limits of agreement run −46.92 to +42.22, a span of 89.14. The mean measurement is around 466, so on any single subject the two instruments can disagree by 19.1% of the quantity being measured.
A correlation measures whether the methods move together, not whether they match. Multiply one method by ten and the correlation is unchanged; the agreement is destroyed.
It also rises with the range of the sample. Measure a wider spread of subjects and r goes up with no change in the instruments at all, which makes it trivially inflatable by choosing who to recruit.
The limits are estimates, and nobody draws their intervals
The limits are printed as though they were known quantities. They are two numbers estimated from the same n subjects as everything else.
On the shipped preset the upper limit is 42.22. Its own 95% confidence interval runs from 25.96 to 58.48.
That is a span of 32.52 on a limit of 42.22 — the limit is uncertain by about three quarters of its own value, at twenty-five subjects.
The standard error of a limit is much larger than that of the bias because it carries the uncertainty in the standard deviation as well as in the mean. That is why a study can pin down the bias and still say almost nothing about the limits.
Bland and Altman recommended at least fifty subjects for this reason, and most published studies use fewer.
When one pair of limits describes nothing
The whole method assumes the difference between the two instruments does not depend on how large the measurement is. When it does, the limits are an average of two different situations.
The proportional-bias preset regresses difference on average and gets a slope of −0.2168 with p below 0.00001. The second method reads about 25% high.
So at the bottom of the range the two nearly agree and at the top they are 30 units apart. A single pair of limits sits between those and describes neither.
The usual fix is to work on the logarithms, which turns a constant ratio into a constant difference and makes the limits interpretable as percentages.
The alternative is to fit the relationship rather than average it away — a Deming or Passing-Bablok regression estimates the slope and intercept between the two methods directly.
Why the x axis is the average and not the reference
It is tempting to plot the difference against the reference method, on the grounds that the reference is the truth. Doing so manufactures a slope out of nothing.
The difference A − B contains B, and so does the x axis. Any measurement error in B appears on both axes with opposite signs, which produces a downward slope even when the two methods agree perfectly on average.
Using the average splits that error between the axes and removes most of the artefact.
It is not a perfect fix — the average still contains both errors — but it is unbiased when the two methods have similar error variances, which is the usual case in a method comparison.
If one method genuinely is a reference standard with negligible error, plotting against it is defensible; the assumption then needs stating, because it is rarely true.
The limits do not say whether the methods are interchangeable
This is the step the statistics cannot take, and the one most often skipped.
Limits of ±44 are fine for some measurements and catastrophic for others. The analysis produces a number; whether it is acceptable is a question about consequences.
The acceptable difference has to be set before the study, from what a wrong reading would cause, not from what the data happened to produce.
Then the comparison is between that limit and the confidence interval on the limits — not against the point estimates, which is why those intervals are printed here.
A study that cannot exclude an unacceptable limit has not shown agreement, even if its point estimates look comfortable.
Two methods cannot agree better than one method agrees with itself
A wide set of limits is often read as a fault in the new method. It may be a fault in the reference.
Every method has its own repeatability. Measure the same subject twice with one instrument and the two readings differ.
The variance of the difference between two methods includes both. So the limits can never be narrower than the repeatability of the worse of the two.
Which means a repeatability study should come first. Run each method twice on the same subjects and put those differences through this same analysis.
If the limits between methods are close to the repeatability limits within a method, the two instruments agree as well as the measurement process allows, and no better instrument would help.
Reporting a Bland-Altman analysis
Five items, and the second and third are the ones that go missing.
Give the bias and the limits, with n. All three, and to enough figures that a reader can recompute the standard deviation.
Give confidence intervals on the limits. Not only on the bias — the limits are the result, and they are the less certain half.
Report the proportional-bias test. A sloping plot invalidates the limits, and eyeballing the scatter is not a test.
State the acceptable difference you set in advance, and say whether the limits clear it.
And do not report a correlation as evidence of agreement. If you report one, say what it is for.
Sources and methodology
References for limits of agreement and their precision.
Method. Differences are taken as A minus B and plotted against the average of the two, never against either method alone — plotting against a supposed reference puts that method’s measurement error on both axes with opposite signs and manufactures a slope out of nothing. The limits use a 1.96 multiplier on the standard deviation of the differences, and each limit carries its own confidence interval built on the larger standard error that accounts for uncertainty in the standard deviation as well as in the mean. Proportional bias is tested by an ordinary regression of difference on average, which is reported with its p-value rather than left to the eye. Pearson’s r is computed on the same pairs and printed, because it is the statistic this analysis exists to displace and the contrast is the point. The suite asserts that the bias is always exactly the difference of the two means, that the limits sit exactly 1.96 standard deviations either side of it, and that swapping the two methods flips the bias and leaves the spread untouched — all on 200 generated studies. That engine is verified on every change against 103 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
Concordance CorrelationLin's concordance correlation with its exact decomposition into precision and accuracy, the scale and location shifts separated, and a bounded interval.
Passing-BablokRobust method-comparison regression with rank-based intervals, the symmetry shift reported, and the pairwise-slope distribution shown.
Deming RegressionRegression when both variables carry error, with a settable error-variance ratio, jackknife intervals and least squares shown as the limiting case.
Intraclass CorrelationAll six ICC forms from one subject-by-rater matrix, with the rater means that drive them apart: one 8x3 matrix gives ICC(1,1) = 0.1277 and ICC(3,1) = 0.9852.
Gauge R&RANOVA gauge repeatability and reproducibility with both acceptance criteria, the operator-by-part interaction tested, and the full variance decomposition.
Correlation CoefficientReports Pearson, Spearman and Kendall together with the scatter plot, and ships Anscombe's quartet built in — four datasets with an identical r of 0.816 that Spearman tells apart.
An educational tool. Limits of agreement describe the data and do not say whether two methods are interchangeable — that requires an acceptable difference set before the study from what a wrong reading would cause. The limits also assume the differences are roughly normal and that the bias does not depend on the measurement level; the proportional-bias test on this page checks the second, and a sloping result invalidates the limits.