Cohen's Kappa
VerifiedReliability (legacy hub)
Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.
Loading teaching datasets…
Or use your own dataset
Loading your datasets…
Agreement between two raters on categorical judgments, corrected for the agreement expected by chance.
Cohen's kappa runs from below 0 (worse than chance) to 1 (perfect). Rough bands: .41–.60 moderate, .61–.80 substantial, > .80 almost perfect.
Worked example
Do two coders agree on classifying 100 responses?
Two coders independently assigned each of 100 responses to one of three categories; kappa corrects their raw agreement for chance.
Agreement was substantial, κ = .73 (raw agreement 82%).
Inter-rater agreement was substantial, Cohen's κ = .73.
Try it yourself: Load this ready-made sample and follow the run above.
When to use it
- Two fixed raters, nominal categoriesThe same two coders classify every case into unordered categories. e.g. two coders assigning 100 responses to one of three themes.
- You need chance-corrected agreementRaw percent agreement flatters raters when one category dominates; kappa subtracts the agreement expected by chance.
When NOT to — use instead
- Three or more ratersKappa is defined for two; generalise to a panel. → Fleiss' kappa
- Ordered categoriesNominal kappa treats a one-grade miss the same as a far miss; weight the disagreements. → Weighted kappa
- Continuous measurementsFor numeric ratings use an agreement coefficient for continuous data. → Intraclass correlation
Assumptions (and what to do if they fail)
Check: When one category is very common, kappa can be low even at 95% agreement. Always report percent agreement and the marginal distribution alongside κ.
If violated: A genuinely reliable coding scheme is dismissed as unreliable purely because the trait is rare.
Check: Cohen's κ assumes the two raters are fixed; a rotating pool of coders needs Fleiss.
If violated: Mixing raters violates the model and the chance-agreement term is mis-estimated.
Check: Each case falls in exactly one category for each rater.
If violated: Overlapping or missing categories make the agreement table ill-defined.
Ready to run a Cohen's Kappa on your own data?
Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.
Run this test →