Inter-rater reliability, with all six forms shown rather than one chosen for you.
All six forms from one matrix
ICC(1,1) = 0.1277 — "unusable". ICC(3,1) = 0.9852 — "excellent". Same eight subjects, same three raters, same numbers. Rater 2 runs about 14 points low and rater 3 about 11 points high, consistently.
8 subjects × 3 raters
ICC(1,1) = 0.127727 … ICC(3,1) = 0.985176
The six forms are not variants of one number. On this matrix the single-measure forms span 0.8574, so naming the form is not a footnote — it is most of the result. F = 200.3810 on 7 and 14 degrees of freedom, p = 6.0940e-13.
Single-measure spread
0.85745
highest minus lowest form
F (subjects)
200.3810
p = 6.094e-13
Mean square between
225.4286
subject-to-subject
Mean square error
1.1250
residual disagreement
All six intraclass correlation forms with what each one assumes
Form
Value
What it assumes
ICC(1,1)
0.127727
Each subject rated by a different, randomly chosen rater
ICC(2,1)
0.323121
Same raters for all subjects, drawn from a wider pool — absolute agreement
ICC(3,1)
0.985176
Same raters, treated as the only ones of interest — consistency
ICC(1,k)
0.305212
As ICC(1,1), but the reported score is the mean of k raters
ICC(2,k)
0.588834
As ICC(2,1), for a mean of k raters
ICC(3,k)
0.995010
As ICC(3,1), for a mean of k raters
Rater means, which is where the spread comes from
Mean score for each rater and its distance from the overall mean
Rater
Mean score
Distance from overall mean
Rater 1
46.62500
+1.1250
Rater 2
32.50000
-13.0000
Rater 3
57.37500
+11.8750
A systematic difference between rater means is disagreement to the absolute-agreement forms (1,1) and (2,1) and invisible to the consistency form (3,1). That single distinction is what produces the spread above.
Report the form, always (…,k) applies only if you average raters in practice Negative values are reported as computed
What this tool shows
One matrix of eight subjects and three raters gives ICC(1,1) = 0.1277 and ICC(3,1) = 0.9852. Unusable or excellent, from the same numbers, depending only on which of the six forms is named. The tool computes all six and prints the rater means that cause the gap, because “the ICC was 0.99” without the form is not a reportable result.
All six ICC forms — (1,1), (2,1), (3,1) and their k-rater counterparts
What each form assumes about the raters, stated next to its value
The spread between the single-measure forms, as a number
Rater means and their distance from the overall mean — the source of the spread
The F test for subject variance, with its degrees of freedom and p-value
Negative ICCs reported as computed rather than floored at zero
Six forms ANOVA mean squares Rater means shown Absolute vs consistency
An ICC reported without its form is not interpretable.
Updated 12 September 2026 · Works in any browser, no installation
An intraclass correlation is the share of total variance that is real subject-to-subject difference rather than measurement disagreement. It is estimated from a one- or two-way ANOVA on a subject-by-rater matrix. The six forms differ in what they treat as error: whether the same raters saw every subject, whether those raters stand in for a wider pool, and whether a rater who is consistently generous counts as disagreeing.
At a glance
Formula shown
From the ANOVA mean squares — MSB between subjects, MSW within, MSR between raters, MSE residual — ICC(1,1) = (MSB − MSW)/(MSB + (k−1)MSW); ICC(2,1) = (MSB − MSE)/(MSB + (k−1)MSE + k(MSR − MSE)/n); ICC(3,1) = (MSB − MSE)/(MSB + (k−1)MSE). The k-rater forms replace the single rating with the mean of k, which is the Spearman–Brown step. Only ICC(2,1) contains the rater term, which is why only it penalises systematic rater bias in the two-way model.
Scenario support
Inter-rater reliability for clinical scales, agreement between measurement devices or laboratories, test-retest reliability across sessions, coding reliability in qualitative and content analysis, judge agreement in competitions and peer review, and any repeated measurement of the same targets where the question is how much of the variation is the target.
Educational estimate
Planning support from the values you enter — not professional advice.
The six forms are six different questions
Shrout and Fleiss named them in 1979 and the naming has been ignored ever since. The forms are not refinements of one another; they answer different questions and can differ by more than the whole usable range.
The first number is the design. 1 means each subject was rated by a different, randomly chosen rater. 2 means the same raters rated everyone and are a sample of a wider pool. 3 means the same raters rated everyone and are the only raters of interest.
The second number is what you will actually use. (…,1) is the reliability of a single rating. (…,k) is the reliability of the mean of k raters, which is always higher, and is the right one only if your real procedure averages raters.
The gap between forms 2 and 3 is the whole question of rater bias. Form 2 counts a rater who is consistently 14 points low as disagreement. Form 3 does not, because it removes the rater effect first.
On the tool’s second preset that single choice moves the answer from 0.1277 to 0.9852. Rater 2 runs about 14 points low and rater 3 about 11 high, consistently across all eight subjects. Every rater ranks the subjects almost identically; none of them agree on the value.
Choosing the form from the study, not from the result
The form is determined by how the data was collected and how the measurement will be used. It is not a modelling preference, and it is certainly not a choice to be made after seeing the values.
Did every subject get the same raters? If not — different clinicians saw different patients — the design is one-way and only form 1 applies. Forms 2 and 3 are not computable from that design even though the arithmetic will produce a number.
Do you care about absolute values or only ranking? If a measurement will be compared against a threshold, systematic bias matters and you want absolute agreement — form 2. If only the ordering matters, consistency — form 3 — is the honest choice.
Will the deployed measurement be one rating or an average? If clinical practice uses a single rater’s score, report (…,1). Reporting (…,k) for a procedure that uses one rater overstates the reliability of what people will actually do.
Reporting the highest of the six is the failure mode this page is about. It is usually ICC(3,k), which is the most permissive combination available, and it is almost never the one the study design supports.
A negative ICC, and why it is not clamped
The ICC is a ratio of variance components, and a variance component estimated by subtraction can come out below zero.
It means the within-subject disagreement exceeds the between-subject variation. Raters disagree about a given subject more than subjects differ from each other — which is worse than no reliability, not equal to it.
Software that returns 0 instead is hiding the diagnostic. A true value of 0 and an estimate of −0.41 mean different things about a dataset, and the second usually means something is wrong upstream.
Small samples produce negative estimates by chance even when the true ICC is positive, which is a reason for a wide confidence interval rather than a reason to round up.
The tool reports the value the arithmetic gives. The third preset is negative on every form, and that is the correct description of raters who disagree more than chance predicts.
ICC, kappa and Pearson answer different things
All three get called “agreement” and only one of them is about agreement in the sense most people mean.
Pearson’s r is blind to systematic bias entirely. A rater who scores every subject exactly 14 points below another correlates with them at r = 1. That is the same blindness ICC(3,1) has, which is why the consistency form and a correlation often land close together.
Pearson also only handles two raters. With three or more you would be averaging pairwise correlations, which is not a defined reliability coefficient. See the correlation coefficient calculator for what it does measure.
For categorical ratings the ICC is the wrong tool.Cohen’s kappa handles nominal categories and corrects for chance agreement, which a variance ratio cannot do.
And for consistency across ITEMS rather than raters, alpha is the counterpart.Cronbach’s alpha is arithmetically close to ICC(3,k) with items in place of raters — the same mean squares, a different substantive question.
The 0.75 and 0.90 thresholds are conventions, not findings
Koo and Li’s bands — below 0.5 poor, 0.5 to 0.75 moderate, 0.75 to 0.90 good, above 0.90 excellent — are widely quoted and were offered as rules of thumb rather than derived cutoffs.
The required reliability depends on the decision. A measurement used to screen a population tolerates far less than one used to decide an individual’s treatment, and no single band covers both.
The confidence interval usually straddles a band boundary. With 20 subjects and 2 raters the interval around an ICC of 0.80 is wide enough to include both “moderate” and “excellent”, which makes the label less informative than the number.
An ICC also depends on the sample’s heterogeneity. The same instrument, with identical measurement error, reports a higher ICC on a more varied sample — because the between-subject variance in the numerator is larger. It is not a property of the instrument alone.
Which makes the standard error of measurement the more portable number. It is in the original units and does not move with the sample’s spread, and reporting it alongside the ICC costs one line.
Reporting an intraclass correlation
Four elements, and the first is the one most commonly omitted.
Name the form. “ICC(2,1) = 0.61” is a result. “ICC = 0.61” is not, because the same data supports five other numbers.
Give the number of subjects and the number of raters. They determine the degrees of freedom and therefore how much the estimate can be trusted.
Give a confidence interval. ICCs from small samples are unstable, and the point estimate alone implies a precision the design does not have.
And say whether the deployed measurement is a single rating or an average. It decides between (…,1) and (…,k), and reporting the wrong one overstates what the procedure achieves in practice.
Sources and methodology
References for the ICC forms and their interpretation.
Method. The four mean squares are computed from the two-way ANOVA decomposition of the subject-by-rater matrix, and all six forms are derived from those same four numbers — so the forms can never disagree about the data underneath them, only about what counts as error. The spread between single-measure forms is computed rather than described, which is what makes the central claim checkable: the preset matrix gives ICC(1,1) = 0.127727 and ICC(3,1) = 0.985176, a spread of 0.857450, with F = 200.3810 and p = 6.09e-13. Negative values are returned as computed and never clamped to zero. Ragged rows, a single subject and a single rater all return no result rather than a number. That engine is verified on every change against 69 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
Cronbach's AlphaAlpha with alpha-if-item-dropped, item-total correlations and the mean inter-item correlation, plus the length table: at a mean inter-item correlation of 0.05, a 100-item scale still reports 0.8403.
Cohen's KappaKappa with the two figures that explain it: the maximum the marginals permit, and PABAK. Two built-in tables with identical 85% agreement give kappas of 0.6995 and 0.3219, and a third with 94.4% agreement gives −0.0234.
Correlation CoefficientReports Pearson, Spearman and Kendall together with the scatter plot, and ships Anscombe's quartet built in — four datasets with an identical r of 0.816 that Spearman tells apart.
One-Way ANOVAThe full F table with eta and omega squared, plus every pairwise gap — because a significant F says something differs and never says which, and ten groups tested pairwise carry a 90% false-positive rate.
VarianceSample and population variance from your data, with a live simulation that shows exactly how much the wrong divisor costs — 20% low at n = 5, closing as the sample grows.
Phi CoefficientPhi for a 2x2 table printed against the ceiling its marginals impose: on [10, 40, 0, 50] every case is exposed, a complete association, and phi is 0.3333333, which is exactly max phi.
An educational tool. The six ICC forms answer different questions and can differ by more than the whole usable range on the same data, so the form must be chosen from the study design rather than from the results — and an ICC depends on how varied the sample is, which means it is not a portable property of the instrument.
Published all six ICC forms from one subject-by-rater matrix, because reporting 'the ICC' without naming the form is the error the page exists to make visible.
Built the central example to settle how much the form matters: one 8x3 matrix gives ICC(1,1) = 0.127727 and ICC(3,1) = 0.985176, a spread of 0.857450. Unusable or excellent, from the same numbers, with F = 200.3810 and p = 6.09e-13 either way.
Printed the rater means beside the forms, since a systematic offset between raters is disagreement to the absolute-agreement forms and invisible to the consistency form — which is the entire source of that spread.
Added a counter-example where the forms agree to within 0.0057, so the reader can tell when the choice is consequential and when it is not.
Reported negative ICCs as computed, and verified the whole set against 69 assertions built on the shared ANOVA decomposition so the six forms can never disagree about the data underneath them.
Add this calculator to your site
Responsive embed — and private: nothing your visitors type leaves their browser.