Weighted Kappa
VerifiedReliability (legacy hub)
Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.
Loading teaching datasets…
Or use your own dataset
Loading your datasets…
Agreement between two raters on ORDERED categories, where near-misses count more than far-misses.
Weighted kappa penalizes disagreements by how far apart they are (linear or quadratic weights) — the right choice when categories are ordinal, e.g. mild / moderate / severe.
Worked example
Do two radiologists agree on a 4-level severity grade?
Two radiologists graded 60 scans on an ordered 4-point scale; quadratic-weighted kappa credits close grades.
Agreement was substantial, weighted κ = .79.
Agreement on the ordinal severity grade was substantial, quadratic-weighted κ = .79.
Try it yourself: Load this ready-made sample and follow the run above.
When to use it
- Two raters, ordered categoriesThe categories have a natural order and a near-miss should count more than a far miss. e.g. two radiologists grading severity mild/moderate/severe/critical.
- You want partial credit for close callsWeighted kappa penalises disagreements by how far apart the grades are, unlike nominal kappa.
When NOT to — use instead
- Categories are unorderedWith no natural order, weighting distances is meaningless. → Cohen's kappa
- Three or more ratersWeighted kappa is a two-rater coefficient. → Fleiss' kappa
- Continuous gradesFor numeric ratings use an intraclass correlation. → Intraclass correlation
Assumptions (and what to do if they fail)
Check: Linear and quadratic weights give different κ; report which you used (quadratic-weighted κ approximates an ICC).
If violated: An unstated scheme makes the value unreproducible and non-comparable across studies.
Check: The weights assume the grades are evenly spaced steps.
If violated: If a one-step gap means more at one end of the scale, the weights mis-score those disagreements.
Check: Skewed grade frequencies can depress κ; report the grade distribution.
If violated: A reliable grading scheme looks poor because most cases fall in one grade.
Ready to run a Weighted Kappa on your own data?
Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.
Run this test →