IRT — Dichotomous

Verified

Psychometrics (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Item Response Theory calibration for binary (0/1) items under three nested models: 1PL/Rasch (item difficulty only, all items equally discriminating), 2PL (difficulty + discrimination), 3PL (adds gues

sing parameter for multiple-choice). Reports per-item difficulty (b) + discrimination (a) + guessing (c) + standard errors, item characteristic curves (ICCs), test information function, person abilities (θ) with conditional SE, and model-comparison LRTs (1PL vs 2PL vs 3PL via -2ΔLL). Engine auto-picks model by information criteria + theoretical considerations.

Worked example

How difficult and discriminating are the test's binary items?

A 2-parameter logistic (2PL) IRT model was fit to 20 right/wrong items (n = 1,000), estimating each item's difficulty and discrimination.

Result

Discriminations ranged 0.8–2.1 (all acceptable) and difficulties spanned −1.9 to +2.2 logits, covering low to high ability; test information peaked near the mean.

How you'd report it (APA)

A 2PL IRT model showed acceptable discriminations (0.8–2.1) and a difficulty range (−1.9 to +2.2 logits) covering the ability continuum.

When to use it

  • High-stakes test calibration (educational testing)
    60-item certification exam (multiple-choice, scored 0/1) on n=2000 examinees.
  • Rasch model — sample-invariant objective measurement
    Math achievement test on n=600 students.

When NOT to — use instead

Assumptions (and what to do if they fail)

Items scored 0 (incorrect) / 1 (correct)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Single underlying latent trait (unidimensionality)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Local independence (item responses independent given θ)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Listwise exclusion of incomplete response setsmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Ready to run a IRT — Dichotomous on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →