Cohen's Kappa

Verified

Reliability (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Agreement between two raters on categorical judgments, corrected for the agreement expected by chance.

Cohen's kappa runs from below 0 (worse than chance) to 1 (perfect). Rough bands: .41–.60 moderate, .61–.80 substantial, > .80 almost perfect.

Worked example

Do two coders agree on classifying 100 responses?

Two coders independently assigned each of 100 responses to one of three categories; kappa corrects their raw agreement for chance.

Result

Agreement was substantial, κ = .73 (raw agreement 82%).

How you'd report it (APA)

Inter-rater agreement was substantial, Cohen's κ = .73.

Try it yourself: Load this ready-made sample and follow the run above.

When to use it

  • Two fixed raters, nominal categories
    The same two coders classify every case into unordered categories. e.g. two coders assigning 100 responses to one of three themes.
  • You need chance-corrected agreement
    Raw percent agreement flatters raters when one category dominates; kappa subtracts the agreement expected by chance.

When NOT to — use instead

  • Three or more raters
    Kappa is defined for two; generalise to a panel. Fleiss' kappa
  • Ordered categories
    Nominal kappa treats a one-grade miss the same as a far miss; weight the disagreements. Weighted kappa
  • Continuous measurements
    For numeric ratings use an agreement coefficient for continuous data. Intraclass correlation

Assumptions (and what to do if they fail)

The kappa paradox with skewed marginalshigh

Check: When one category is very common, kappa can be low even at 95% agreement. Always report percent agreement and the marginal distribution alongside κ.

If violated: A genuinely reliable coding scheme is dismissed as unreliable purely because the trait is rare.

Same two raters across all casesmedium

Check: Cohen's κ assumes the two raters are fixed; a rotating pool of coders needs Fleiss.

If violated: Mixing raters violates the model and the chance-agreement term is mis-estimated.

Categories are mutually exclusive and exhaustivemedium

Check: Each case falls in exactly one category for each rater.

If violated: Overlapping or missing categories make the agreement table ill-defined.

Ready to run a Cohen's Kappa on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →