Aiken's V
VerifiedReliability (legacy hub)
Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.
Loading teaching datasets…
Or use your own dataset
Loading your datasets…
Content-validity agreement — Aiken's V summarizes how strongly an expert panel endorses each item's relevance.
Aiken's V rescales expert ratings to 0–1 with a significance test; higher V means stronger consensus that an item measures the intended content.
Worked example
Do experts agree the items are relevant?
Eight experts rated each item's relevance on a 1–5 scale; Aiken's V quantifies content-validity agreement per item.
12 of 15 items reached Aiken's V ≥ .80, supporting content validity (the three weakest items fell below and would be revised).
Content validity was supported, with 12 of 15 items at Aiken's V ≥ .80.
Try it yourself: Load this ready-made sample and follow the run above.
When to use it
- Quantifying expert content-validity ratingsA panel rates each item's relevance on an ordered scale (e.g. 1–5) and you want a per-item index with a significance test. e.g. 8 experts rating 15 items.
- You need item-level decisionsV is computed per item, so it tells you which specific items to keep, revise, or drop.
When NOT to — use instead
- Experts rate only relevant / not relevantA binary relevance judgement is the CVI's territory, not a rescaled rating. → Content Validity Index
- You want the 'essential' proportionLawshe's essential/useful/not-necessary format gives the CVR. → Content Validity Ratio
Assumptions (and what to do if they fail)
Check: The rating categories (1–5, say) are equally spaced and the number of scale points is stated — V's formula depends on the scale minimum and range.
If violated: A wrong scale range rescales V incorrectly; a value can even exceed its own 0–1 bound (a tell that the scale was mis-set).
Check: Content-validity panels are typically 5–10 experts; too few makes V and its significance unstable.
If violated: A single expert swings the index; the significance test has almost no power.
Check: Ratings are made without consultation between experts.
If violated: Correlated ratings overstate consensus — the agreement is partly shared opinion, not independent endorsement.
Ready to run a Aiken's V on your own data?
Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.
Run this test →