Differential Item Func.

Coming soon

Psychometrics (legacy hub)

This test is implemented and is currently going through StatMinds’ production verification: every statistic is independently checked against a trusted reference (scipy / R), locked with regression tests, and the screen is exercised across assumption-met/violated and significant/non-significant scenarios before it opens up.

See what’s live now

Detects items that function DIFFERENTLY across reference and focal groups (e.g., gender, race, language) after controlling for ability.

Two complementary approaches: (1) Mantel-Haenszel (MH) — chi-square-based, gives MH common odds ratio + ETS A/B/C effect-size classification; (2) logistic regression DIF — fits logit(P_correct) on ability + group + group×ability, R² change quantifying uniform vs non-uniform DIF. Reports per-item MH α, ETS classification, logistic R²-change, and DIF-flagged items for revision or removal.

Worked example

Does an item behave differently for men and women at the same ability?

Differential item functioning was tested across sex, comparing item responses at matched trait levels (Mantel-Haenszel / IRT).

Result

One of 20 items showed significant uniform DIF (ΔMH = 1.9, p = .004); the rest were DIF-free, so the scale is largely fair across sex.

How you'd report it (APA)

DIF analysis flagged one item with significant uniform DIF across sex (p = .004); the remaining items functioned equivalently.

When to use it

  • Test fairness audit — bias detection by group
    100-item certification exam screened for DIF by ethnicity (n=400 reference White, n=200 focal Black).
  • Cross-language scale validation
    20-item depression scale translated English → Spanish, administered to bilingual sample (n=400 English, n=380 Spanish).

When NOT to — use instead

Assumptions (and what to do if they fail)

Group variable is exactly 2 levels (reference vs focal)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Total score is an appropriate matching variable (unidimensionality)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Sample size ≥ 200 per group preferredmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Listwise exclusion of missing item responsesmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Ready to run a Differential Item Func. on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →