IRT — Dichotomous
VerifiedPsychometrics (legacy hub)
Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.
Loading teaching datasets…
Or use your own dataset
Loading your datasets…
Item Response Theory calibration for binary (0/1) items under three nested models: 1PL/Rasch (item difficulty only, all items equally discriminating), 2PL (difficulty + discrimination), 3PL (adds gues
sing parameter for multiple-choice). Reports per-item difficulty (b) + discrimination (a) + guessing (c) + standard errors, item characteristic curves (ICCs), test information function, person abilities (θ) with conditional SE, and model-comparison LRTs (1PL vs 2PL vs 3PL via -2ΔLL). Engine auto-picks model by information criteria + theoretical considerations.
Worked example
How difficult and discriminating are the test's binary items?
A 2-parameter logistic (2PL) IRT model was fit to 20 right/wrong items (n = 1,000), estimating each item's difficulty and discrimination.
Discriminations ranged 0.8–2.1 (all acceptable) and difficulties spanned −1.9 to +2.2 logits, covering low to high ability; test information peaked near the mean.
A 2PL IRT model showed acceptable discriminations (0.8–2.1) and a difficulty range (−1.9 to +2.2 logits) covering the ability continuum.
When to use it
- High-stakes test calibration (educational testing)60-item certification exam (multiple-choice, scored 0/1) on n=2000 examinees.
- Rasch model — sample-invariant objective measurementMath achievement test on n=600 students.
When NOT to — use instead
- Polytomous / Likert items1PL/2PL/3PL are for binary items. → IRT \u2014 Polytomous
- Sample too small (n < 200)IRT calibration needs adequate sample. → Split-Half (+ KR-20/KR-21)
- Multi-dimensional test (multiple constructs)Standard IRT assumes unidimensionality. → Exploratory FA
- Differential item functioning (DIF) detectionDIF compares item parameters across groups — separate analysis from baseline calibration. → Differential Item Func.
Assumptions (and what to do if they fail)
Check: See the assumption diagnostics in the workspace.
If violated: The workspace flags this and suggests a robust or nonparametric alternative.
Check: See the assumption diagnostics in the workspace.
If violated: The workspace flags this and suggests a robust or nonparametric alternative.
Check: See the assumption diagnostics in the workspace.
If violated: The workspace flags this and suggests a robust or nonparametric alternative.
Check: See the assumption diagnostics in the workspace.
If violated: The workspace flags this and suggests a robust or nonparametric alternative.
Ready to run a IRT — Dichotomous on your own data?
Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.
Run this test →