Twelve manuscripts, five reviewers each, three verdicts. κ is 0.3577 — “fair” — with p = 3.4×10⁻⁸. That tiny p-value tests whether κ is zero, and says nothing about how precisely κ is pinned down: the bootstrap interval over manuscripts runs from 0.1093 to 0.5392, which spans “slight” to “moderate”. Look at the per-category table: reviewers agree on rejection (0.4750) and barely at all on revision (0.1892).
12 subjects · 5 raters each · 3 categories
κ = 0.3577 (fair)
Raters agreed on 57.50% of rater pairs against 33.83% expected from the marginal rates alone. Testing κ = 0 gives z = 5.520, p = 3.39e-8. A bootstrap over subjects puts κ between 0.1093 and 0.5392 — the p-value says κ is not zero, and that interval says how little else is settled.
Fleiss' κ
0.3577
fair
Bootstrap 95% interval
0.109 to 0.539
over 12 subjects
Observed agreement
57.50%
expected 33.83%
Unanimous subjects
16.7%
2 of 12
Agreement one category at a time
Category-specific kappa and how often each category was used
Category
Share of ratings
κ for this category
Against overall
Reject
33.33%
0.4750
+0.1173
Revise
38.33%
0.1892
-0.1685
Accept
28.33%
0.4254
+0.0678
A single overall κ averages over categories that can behave completely differently. Here the spread runs from 0.1892 to 0.4750, and the low one is usually the category worth rewriting in the coding manual.
How the two agreements are built
Observed and expected agreement and the kappa built from them
Quantity
Value
What it is
P̄ (observed)
0.575000
share of rater pairs who agreed, averaged over subjects
P̄ₑ (expected)
0.338333
sum of squared category rates — agreement from the margins alone
κ
0.357683
(P̄ − P̄ₑ) / (1 − P̄ₑ): the share of the available room that was used
SE under κ = 0
0.064797
for the test only — not a width for the estimate
Any number of raters Per-category breakdown Bootstrap interval, not just a test Nominal categories only
What this tool shows
The twelve-manuscript preset gives κ = 0.3577 with p = 3.4×10⁻⁸, and a bootstrap interval of 0.1093 to 0.5392. The p-value tests whether κ is zero; it says nothing about how precisely κ is pinned down, and that interval spans “slight” to “moderate”. The per-category table on the same data runs from 0.1892 to 0.4750 — a single overall number hides which category the raters are actually failing on.
Fleiss’ kappa for any fixed number of raters per subject, with different raters allowed per subject
A category-by-category kappa, which is usually where the problem actually is
A bootstrap interval over subjects — the null standard error cannot supply one
Observed and expected agreement reported separately, not folded into one number
The share of subjects on which every rater agreed
The Landis and Koch verbal band, with the reasons not to lean on it
Any rater count Per category Bootstrap interval Both agreements
Nominal categories only. Ordered ratings want a weighted kappa.
Updated 13 September 2026 · Works in any browser, no installation
Fleiss’ kappa is the share of the available agreement that the raters actually achieved. It compares how often pairs of raters put a subject in the same category against how often they would by chance, given how often each category gets used at all. Unlike Cohen’s kappa it handles any number of raters, and it does not require the same people to rate every subject — only that the number of raters per subject is constant.
At a glance
Formula shown
For each subject i, Pᵢ = (Σⱼnᵢⱼ² − n) / (n(n−1)) is the proportion of rater pairs who agreed, with n raters per subject. P̄ is the mean of those. Each category’s overall rate is pⱼ = Σᵢnᵢⱼ / (N·n), and the chance agreement is P̄ₑ = Σⱼpⱼ². Then κ = (P̄ − P̄ₑ) / (1 − P̄ₑ). The denominator is the room left above chance, which is why a kappa is small whenever one category dominates the margins — there is very little room to beat.
Scenario support
Content analysis and qualitative coding with three or more coders, diagnostic agreement across a panel of clinicians, manuscript or grant screening, annotation quality for a machine-learning dataset, and any pilot round of coding where the question is whether the manual is clear enough to use.
Educational estimate
Planning support from the values you enter — not professional advice.
p = 0.000000034 and an interval from 0.11 to 0.54
Almost every Fleiss’ kappa calculator reports a standard error. Almost none of them say what it is for, and it is not for the interval most people build with it.
On the shipped preset κ = 0.3577 and p = 3.4×10⁻⁸. That is overwhelming evidence against κ = 0.
The standard error behind it, 0.0648, is derived under the assumption that κ IS zero. It is a null standard error. Using it to draw an interval around 0.3577 assumes the thing the test just rejected.
The bootstrap interval over subjects runs from 0.1093 to 0.5392 — 1.69× as wide as the null standard error implies, and it spans two of the Landis and Koch bands.
So the honest reading is: agreement is clearly better than chance, and twelve subjects cannot tell you much more than that. Both numbers are printed here for exactly that reason.
The overall κ hides which category is broken
A single agreement figure is an average over categories that often behave nothing like each other, and the average is the least actionable summary available.
On the peer-review preset the overall κ is 0.3577. The category kappas are 0.4750 for reject, 0.1892 for revise and 0.4254 for accept.
Reviewers agree tolerably on rejecting and on accepting, and barely at all on revising. That is a specific, fixable finding; “fair agreement” is not.
A low category kappa usually means the definition is doing too much work. “Revise” is a residual category here — the one chosen when neither of the others clearly applies — and residual categories reliably have the worst agreement.
Which is where a coding manual gets rewritten, and why the per-category table is the first place to look after the headline number.
Why not just report percent agreement
Raw agreement is easy to explain and easy to inflate, and the fourth preset shows exactly how.
That preset has an observed agreement of 0.333333 — a third of all rater pairs agree. Reported as “33% agreement”, it sounds like a number.
Expected agreement on the same margins is also exactly 0.333333, so κ is exactly zero. Every category kappa is zero too. The raters have achieved nothing.
The effect runs the other way as well. If one category is used 90% of the time, raters agree constantly by accident, and percent agreement looks excellent while κ stays near zero.
This is the “kappa paradox” — high agreement with a low kappa — and it is not a flaw in kappa. It is kappa correctly reporting that skewed margins leave almost no room to beat chance.
Fleiss, Cohen and the intraclass correlation
Three agreement statistics get used interchangeably and are not interchangeable.
Cohen’s kappa is for exactly two raters who both rate every subject, and it can be weighted for ordered categories.
Fleiss’ kappa is for a fixed number of raters per subject, who need not be the same people each time — which is what makes it right for crowd annotation and for panels where reviewers rotate.
An intraclass correlation is for numeric ratings rather than categories, and it distinguishes consistency from absolute agreement.
Fleiss’ kappa treats every disagreement as equally bad. If your categories are ordered — mild, moderate, severe — confusing mild with severe should cost more than confusing mild with moderate, and only a weighted statistic will say so.
Checked against the paper that introduced it
The second preset is not an illustration. It is the worked example from Fleiss’ 1971 paper, and it is on this page as a check anyone can repeat.
Thirty patients, six psychiatrists, five diagnostic categories. The paper reports κ = 0.430.
This calculator returns 0.4302 from the table, with observed agreement 0.5556 and expected agreement 0.2199.
The five category kappas match too: 0.2448, 0.2448, 0.5200, 0.4711 and 0.5661 against the paper’s 0.245, 0.245, 0.520, 0.471 and 0.566.
None of those figures is hard-coded. They are computed from the counts in the textarea, which means editing a single cell changes them — and the verification suite asserts all six against the published values on every build.
The Landis and Koch bands are a convention
The familiar labels — slight, fair, moderate, substantial, almost perfect — come from a 1977 paper that describes them as arbitrary.
They have no theoretical basis and no field-specific calibration. A κ of 0.6 may be excellent for open-ended qualitative coding and unacceptable for a diagnostic decision.
They ignore the number of categories. Agreement on two categories and on twelve are very different achievements at the same κ.
They ignore the interval. The preset here has a point estimate in “fair” and an interval covering “slight” through “moderate”.
Set the threshold before collecting the ratings, based on what the ratings will be used for — which also removes the temptation to pick a band after seeing the number.
Reporting a Fleiss’ kappa
Five items. The first three are the ones that let a reader judge the rest.
Give the number of subjects, raters and categories. κ means different things at different table shapes.
Give the observed agreement alongside κ. Together they show whether a low κ comes from poor agreement or from skewed margins.
Give an interval, and say how it was produced. A bootstrap over subjects is honest; the null standard error used as a half-width is not.
Give the per-category kappas when there are more than two categories. It is where the actionable information lives.
And state the threshold you set in advance, rather than the band the number happened to land in.
Sources and methodology
References for Fleiss’ kappa and the agreement literature.
Method. κ is computed from the subject-by-category count table exactly as Fleiss defined it, and the 1971 paper’s own worked example ships as a preset so the implementation can be checked against the source rather than against another calculator: the suite asserts κ = 0.430, both agreement figures and all five category-specific kappas against the published values. The standard error is the Fleiss asymptotic variance under the null that raters assign categories independently at the observed marginal rates, and it is labelled as a test statistic rather than used as an interval half-width. The interval instead comes from 1,200 bootstrap resamples of the subjects, with a fixed seed so the figure is reproducible; subjects are the independent unit, and resampling ratings instead would understate the uncertainty. The suite also asserts that κ is unchanged by the order of the subjects and by the order of the categories, that unanimous ratings give exactly 1, and that ragged rows, fractional counts and a degenerate single-category margin all return no result rather than a plausible-looking number. That engine is verified on every change against 134 assertions. The count and the per-case breakdown are published on the formula verification page.
Related calculators
Where this goes next:
Cohen's KappaKappa with the two figures that explain it: the maximum the marginals permit, and PABAK. Two built-in tables with identical 85% agreement give kappas of 0.6995 and 0.3219, and a third with 94.4% agreement gives −0.0234.
Intraclass CorrelationAll six ICC forms from one subject-by-rater matrix, with the rater means that drive them apart: one 8x3 matrix gives ICC(1,1) = 0.1277 and ICC(3,1) = 0.9852.
Cronbach's AlphaAlpha with alpha-if-item-dropped, item-total correlations and the mean inter-item correlation, plus the length table: at a mean inter-item correlation of 0.05, a 100-item scale still reports 0.8403.
Cramer's VThe effect size a chi-square test does not give you, with your own table rescaled six ways: the same 3x3 pattern gives V = 0.27136852 at n = 90 and at n = 1,800 while p falls from 0.0101 to 3.6e-56.
Chi-SquareGoodness of fit and tests of independence with every expected count and per-cell contribution shown — because the validity condition is about expected counts, not observed ones, and most calculators hide them.
Cliff's DeltaNon-parametric effect size from the full pairwise win–loss–tie count, with a DeLong interval, magnitude bands and Cohen's d for comparison.
An educational tool. Fleiss’ kappa treats every disagreement as equally serious, so ordered categories — mild, moderate, severe — are better served by a weighted kappa. It is also sensitive to the marginal rates: when one category dominates, agreement can be high and kappa near zero at the same time, which is a correct report rather than a fault.