Item Text Quality Audit

Verified

Psychometrics (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Text-level audit of SCALE ITEMS for linguistic and psychometric quality concerns.

Six checks per item: (1) READABILITY (Flesch-Kincaid grade level, SMOG index); (2) DOUBLE-BARRELED detection (two claims joined by 'and' / 'or'); (3) NEGATION (items containing 'not', 'no', 'never' — often harder to process); (4) ABSOLUTE language ('always' / 'never' — often reduces variance); (5) LENGTH (words + characters); (6) SEMANTIC SIMILARITY (via sentence embeddings — flags near-duplicate items). Per-item flag list + overall scale quality summary. Pair with content validity (expert ratings) and CTT item analysis (psychometric behavior) for complete item-level audit.

Worked example

Are any items problematic overall?

An item-quality review combines difficulty, discrimination and reliability-if-deleted into a per-item flag.

Result

17 of 20 items were healthy; three were flagged (low discrimination or redundancy) for revision or removal.

How you'd report it (APA)

Item-quality screening flagged three of 20 items (low discrimination / redundancy) for revision.

When to use it

  • New scale linguistic quality assurance
    30-item depression scale draft: Flesch-Kincaid grade 9.4 (target: ≤ 8 for general population).
  • Scale revision after psychometric issues traced to text
    Pilot flags item 14 as low-discrimination (item-total r = 0.15).

When NOT to — use instead

  • Respondent-level data quality
    Item-quality audits ITEM TEXT. Survey Data Quality Screen
  • Construct validity from expert ratings
    Item-quality is linguistic. Content Validity
  • Psychometric item behaviour (difficulty, discrimination)
    For item psychometrics (CTT / IRT statistics) use ctt_item_analysis or irt_dichotomous / irt_polytomous. CTT Item Analysis
  • Translation quality between languages
    For cross-language item quality use cross_cultural (source-vs-target + DIF). Cross-Cultural Validity

Assumptions (and what to do if they fail)

Item texts are in English (for readability metrics)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Each line corresponds to one item's full promptmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

≥ 3 items for similarity-matrix screeningmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Ready to run a Item Text Quality Audit on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →