KR-20 / KR-21

Verified

Reliability (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Internal-consistency reliability for a test of right/wrong items — the dichotomous-item version of Cronbach's alpha.

Kuder-Richardson 20 estimates how consistently a set of 0/1 items measures one ability. Ranges 0–1; ≥ .70 is acceptable.

Worked example

Is a 25-item multiple-choice test internally consistent?

500 students' right/wrong responses to 25 items were analysed with KR-20.

Result

KR-20 = .87 — good internal consistency for the ability test.

How you'd report it (APA)

The 25-item test showed good internal consistency, KR-20 = .87.

Try it yourself: Load this ready-made sample and follow the run above.

When to use it

  • Right/wrong test items
    A set of dichotomous (0/1) items scored for one ability. e.g. a 25-item multiple-choice test.
  • Internal-consistency reliability
    KR-20 is Cronbach's alpha specialised to binary items — one number for how consistently the items measure the same thing.

When NOT to — use instead

  • Likert / polytomous items
    For graded responses use the continuous-item reliability coefficient. Alpha & Omega
  • Items with unequal loadings you want to respect
    For a model-based reliability that does not assume tau-equivalence, use omega/composite reliability. Composite reliability

Assumptions (and what to do if they fail)

Items are dichotomous and unidimensionalhigh

Check: Every item is scored 0/1 and the items measure a single ability.

If violated: KR-20 on a multidimensional test underestimates reliability and misrepresents what one score means.

Tau-equivalence (equal true-score loadings)medium

Check: Like alpha, KR-20 assumes items contribute equally; unequal items make it a lower bound.

If violated: Reliability is understated — the test may be better than KR-20 suggests.

Reliability is sample- and length-dependentlow

Check: KR-20 rises with more items and with a spread of ability in the sample.

If violated: A low value in a homogeneous sample, or on a short test, is not necessarily a bad test.

Ready to run a KR-20 / KR-21 on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →