Fleiss' Kappa
VerifiedReliability (legacy hub)
Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.
Loading teaching datasets…
Or use your own dataset
Loading your datasets…
Agreement among three or more raters on categorical judgments, corrected for chance.
Fleiss' kappa extends Cohen's kappa to any number of raters (who need not be the same across subjects), with the same interpretive bands.
Worked example
Do five reviewers agree when triaging manuscripts?
Five reviewers each rated 80 manuscripts into three decisions; Fleiss' kappa summarizes multi-rater agreement.
Agreement was moderate, κ = .48.
Agreement among the five reviewers was moderate, Fleiss' κ = .48.
Try it yourself: Load this ready-made sample and follow the run above.
When to use it
- Three or more raters, nominal categoriesA panel classifies each case; the raters need not be the same individuals across cases. e.g. 5 reviewers triaging 80 manuscripts.
- Fixed number of ratings per caseEach subject receives the same number of ratings, which Fleiss' model requires.
When NOT to — use instead
- Exactly two ratersUse the two-rater coefficient. → Cohen's kappa
- Ordered categoriesFleiss treats all disagreements alike; for ordinal grades weight them (with two raters). → Weighted kappa
- Continuous ratingsFor numeric scores use an intraclass correlation. → Intraclass correlation
Assumptions (and what to do if they fail)
Check: Fleiss' κ assumes a fixed count of ratings for every case; unequal counts break the formula.
If violated: The chance-agreement term is wrong and κ is not interpretable — balance the design or use a model that allows missingness.
Check: As with Cohen's, a dominant category can depress κ despite high agreement; report the category proportions.
If violated: A reliable scheme is judged unreliable because the trait is rare.
Check: Each rater places each case in exactly one category.
If violated: Overlapping categories make the rating counts ambiguous.
Ready to run a Fleiss' Kappa on your own data?
Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.
Run this test →