Split-Half (+ KR-20/KR-21)

Verified

Psychometrics (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Splits a k-item scale into two halves (odd-even, random, or first-second), correlates the two halves, and applies the Spearman-Brown prophecy formula to estimate full-scale reliability.

Reports Spearman-Brown corrected coefficient, Guttman's λ₄ (best of all possible splits), and KR-20 / KR-21 specifically for dichotomous (0/1) items. The historical reliability tool for binary-item achievement tests; remains useful when α / ω require assumptions the data won't support.

Worked example

Do two halves of the test agree?

Items were split into two halves; the correlation between half-scores, Spearman-Brown corrected, estimates reliability.

Result

Split-half reliability was strong, Spearman-Brown corrected r = .85.

How you'd report it (APA)

Split-half reliability (Spearman-Brown corrected) was .85.

When to use it

  • Binary-item achievement / knowledge test (KR-20)
    40-item multiple-choice biology exam administered to 300 students.
  • Short scale where α may be unstable — split-half companion
    5-item burnout scale with α = 0.71 (CI 0.62, 0.79).

When NOT to — use instead

  • Likert / continuous items with stable factor structure
    Use Cronbach's α (general) or McDonald's ω (modern) — they are more efficient than split-half for continuous items. Cronbach's Alpha
  • Inter-rater agreement
    Split-half is internal-consistency, not rater agreement. Inter-Rater Reliability
  • Multi-dimensional scale
    Compute split-half / KR-20 PER subscale, not on the full instrument. Split-Half (+ KR-20/KR-21)
  • Item Response Theory analysis is feasible
    When n is adequate (≥ 200) for IRT, 2PL / 3PL gives more information per item than KR-20. IRT \u2014 Dichotomous

Assumptions (and what to do if they fail)

Items can be meaningfully combined into a single compositemedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

The two halves are approximately parallel (similar mean + variance)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

For KR-20 / KR-21: items are dichotomous (0/1)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Listwise exclusion of incomplete response setsmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Ready to run a Split-Half (+ KR-20/KR-21) on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →