Weighted Kappa

Verified

Reliability (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Agreement between two raters on ORDERED categories, where near-misses count more than far-misses.

Weighted kappa penalizes disagreements by how far apart they are (linear or quadratic weights) — the right choice when categories are ordinal, e.g. mild / moderate / severe.

Worked example

Do two radiologists agree on a 4-level severity grade?

Two radiologists graded 60 scans on an ordered 4-point scale; quadratic-weighted kappa credits close grades.

Result

Agreement was substantial, weighted κ = .79.

How you'd report it (APA)

Agreement on the ordinal severity grade was substantial, quadratic-weighted κ = .79.

Try it yourself: Load this ready-made sample and follow the run above.

When to use it

  • Two raters, ordered categories
    The categories have a natural order and a near-miss should count more than a far miss. e.g. two radiologists grading severity mild/moderate/severe/critical.
  • You want partial credit for close calls
    Weighted kappa penalises disagreements by how far apart the grades are, unlike nominal kappa.

When NOT to — use instead

Assumptions (and what to do if they fail)

State the weighting schemehigh

Check: Linear and quadratic weights give different κ; report which you used (quadratic-weighted κ approximates an ICC).

If violated: An unstated scheme makes the value unreproducible and non-comparable across studies.

Equal-interval ordered categoriesmedium

Check: The weights assume the grades are evenly spaced steps.

If violated: If a one-step gap means more at one end of the scale, the weights mis-score those disagreements.

The kappa paradox can still bitemedium

Check: Skewed grade frequencies can depress κ; report the grade distribution.

If violated: A reliable grading scheme looks poor because most cases fall in one grade.

Ready to run a Weighted Kappa on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →