Content Validity

Verified

Psychometrics (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Quantifies whether a scale's items adequately represent the construct domain, judged by a panel of CONTENT EXPERTS (typically 5-10) rating each item's relevance and clarity.

Reports four complementary statistics: (1) Lawshe's CVR (Content Validity Ratio) per item; (2) Polit-Beck I-CVI (item-level CVI) and S-CVI (scale-level CVI, both /Ave and /UA versions); (3) Aiken's V (proportion of agreement above chance); (4) Fleiss' κ across raters (chance-corrected agreement). Items with CVR / I-CVI below thresholds are flagged for revision or removal BEFORE pilot testing.

Worked example

Do expert ratings support the items' content validity?

Ten experts rated each item's relevance; the item-level Content Validity Index (I-CVI) and scale-level S-CVI summarise agreement.

Result

18 of 20 items had I-CVI ≥ .78 (acceptable) and the scale-level S-CVI/Ave was .91, supporting content validity.

How you'd report it (APA)

Content validity was supported, S-CVI/Ave = .91, with 18 of 20 items at I-CVI ≥ .78.

When to use it

  • Expert-panel content validation (pre-pilot)
    30-item depression scale candidate, 8 expert raters.
  • Re-content-validation after scale revision
    Pilot of 25-item scale flags items 8 + 14 + 22 as psychometrically weak (item-total r < 0.20).

When NOT to — use instead

Assumptions (and what to do if they fail)

Rows are independent expert ratersmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Columns are individual itemsmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Ratings are 1–4 (relevance scale)medium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Listwise exclusion of missing ratings per itemmedium

Check: See the assumption diagnostics in the workspace.

If violated: The workspace flags this and suggests a robust or nonparametric alternative.

Ready to run a Content Validity on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →