CTT Item Analysis

Verified

Psychometrics (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Per-item descriptive psychometrics for a unidimensional or subscale-decomposed instrument.

Reports per item: difficulty (mean / proportion correct), discrimination (item-total corrected r, with α-if-deleted), distractor analysis (for multiple-choice tests), floor / ceiling proportion, missing-data rate, and reviewer flags (low discrimination, extreme difficulty, high missingness, redundancy). The standard first-pass diagnostic before reliability or factor analysis — identifies items that are weak / redundant / mis-keyed and should be revised or dropped.

Worked example

Which items are pulling their weight?

Classical item analysis reports each item's difficulty (p) and corrected item-total correlation (discrimination) for a 25-item test.

Result

Most items discriminated well (item-total r .30–.55); two items with r < .20 and one very easy item (p = .96) were flagged for revision.

How you'd report it (APA)

Classical item analysis flagged three weak items (item-total r < .20 or p = .96); the rest discriminated adequately.

When to use it

  • Scale-validation pre-screening (before reliability / EFA)
    30-item adapted depression scale piloted on 400 patients.
  • Achievement test — distractor analysis on multiple-choice items
    100-item multiple-choice biology test.

When NOT to — use instead

  • Item-Response-Theory (IRT) is feasible (n ≥ 300)
    IRT (1PL/2PL/3PL or GRM/PCM) gives more information per item than CTT — item characteristic curves, person-item targeting, etc. IRT \u2014 Dichotomous
  • Scale already validated; just need composite reliability
    CTT item analysis is for pre-validation screening. Cronbach's Alpha
  • Single item without subscale context
    Item-total correlations require a composite to correlate against. Survey Data Quality Screen
  • Inter-rater agreement on items
    CTT item analysis is for rater-agnostic item statistics. Inter-Rater Reliability

Assumptions (and what to do if they fail)

Items share a common construct (within each reported subscale)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Difficulty interpretation: p ≈ 0.30–0.70 optimal for maximum discriminationmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Discrimination interpretation: r_item-rest ≥ .30 (Kline 2015)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Listwise exclusion of incomplete response setsmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Ready to run a CTT Item Analysis on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →