Fleiss' Kappa

Verified

Reliability (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Agreement among three or more raters on categorical judgments, corrected for chance.

Fleiss' kappa extends Cohen's kappa to any number of raters (who need not be the same across subjects), with the same interpretive bands.

Worked example

Do five reviewers agree when triaging manuscripts?

Five reviewers each rated 80 manuscripts into three decisions; Fleiss' kappa summarizes multi-rater agreement.

Result

Agreement was moderate, κ = .48.

How you'd report it (APA)

Agreement among the five reviewers was moderate, Fleiss' κ = .48.

Try it yourself: Load this ready-made sample and follow the run above.

When to use it

  • Three or more raters, nominal categories
    A panel classifies each case; the raters need not be the same individuals across cases. e.g. 5 reviewers triaging 80 manuscripts.
  • Fixed number of ratings per case
    Each subject receives the same number of ratings, which Fleiss' model requires.

When NOT to — use instead

Assumptions (and what to do if they fail)

Same number of ratings per subjecthigh

Check: Fleiss' κ assumes a fixed count of ratings for every case; unequal counts break the formula.

If violated: The chance-agreement term is wrong and κ is not interpretable — balance the design or use a model that allows missingness.

The kappa paradox still appliesmedium

Check: As with Cohen's, a dominant category can depress κ despite high agreement; report the category proportions.

If violated: A reliable scheme is judged unreliable because the trait is rare.

Categories mutually exclusivemedium

Check: Each rater places each case in exactly one category.

If violated: Overlapping categories make the rating counts ambiguous.

Ready to run a Fleiss' Kappa on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →