Differential Item Func.
Coming soonPsychometrics (legacy hub)
This test is implemented and is currently going through StatMinds’ production verification: every statistic is independently checked against a trusted reference (scipy / R), locked with regression tests, and the screen is exercised across assumption-met/violated and significant/non-significant scenarios before it opens up.
See what’s live nowDetects items that function DIFFERENTLY across reference and focal groups (e.g., gender, race, language) after controlling for ability.
Two complementary approaches: (1) Mantel-Haenszel (MH) — chi-square-based, gives MH common odds ratio + ETS A/B/C effect-size classification; (2) logistic regression DIF — fits logit(P_correct) on ability + group + group×ability, R² change quantifying uniform vs non-uniform DIF. Reports per-item MH α, ETS classification, logistic R²-change, and DIF-flagged items for revision or removal.
Worked example
Does an item behave differently for men and women at the same ability?
Differential item functioning was tested across sex, comparing item responses at matched trait levels (Mantel-Haenszel / IRT).
One of 20 items showed significant uniform DIF (ΔMH = 1.9, p = .004); the rest were DIF-free, so the scale is largely fair across sex.
DIF analysis flagged one item with significant uniform DIF across sex (p = .004); the remaining items functioned equivalently.
When to use it
- Test fairness audit — bias detection by group100-item certification exam screened for DIF by ethnicity (n=400 reference White, n=200 focal Black).
- Cross-language scale validation20-item depression scale translated English → Spanish, administered to bilingual sample (n=400 English, n=380 Spanish).
When NOT to — use instead
- Single group — no comparison groupDIF requires ≥ 2 groups to compare. → IRT \u2014 Dichotomous
- Items not yet calibratedDIF needs total-score or IRT-θ as the matching variable. → IRT \u2014 Dichotomous
- Tiny focal-group sample (n_focal < 100)DIF analysis needs adequate focal-group sample. → CTT Item Analysis
- Construct-level invariance questionDIF is item-level. → Measurement Invariance
Assumptions (and what to do if they fail)
Check: See the assumption diagnostics in the workspace.
If violated: The workspace flags this and suggests a robust or nonparametric alternative.
Check: See the assumption diagnostics in the workspace.
If violated: The workspace flags this and suggests a robust or nonparametric alternative.
Check: See the assumption diagnostics in the workspace.
If violated: The workspace flags this and suggests a robust or nonparametric alternative.
Check: See the assumption diagnostics in the workspace.
If violated: The workspace flags this and suggests a robust or nonparametric alternative.
Ready to run a Differential Item Func. on your own data?
Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.
Run this test →