Intraclass Correlation (ICC)
VerifiedReliability (legacy hub)
Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.
Loading teaching datasets…
Or use your own dataset
Loading your datasets…
Intraclass correlation — agreement or consistency among repeated measurements (multiple raters, or the same measure over time).
The ICC is the share of total variance due to real differences between subjects rather than measurement error. Different forms (ICC(1), (2,1), (3,k)) suit rater vs test-retest designs; ≥ .75 is good.
Worked example
Do three clinicians agree when rating the same patients?
Three raters scored 40 patients on a continuous scale; a two-way ICC quantifies their agreement.
Agreement was good, ICC(2,1) = .76, 95% CI [.64, .86].
Inter-rater agreement was good, ICC(2,1) = .76, 95% CI [.64, .86].
Try it yourself: Load this ready-made sample and follow the run above.
When to use it
- Agreement on continuous measurementsSeveral raters score the same subjects on a numeric scale, or the same instrument is repeated over time. e.g. 3 clinicians rating 40 patients.
- Test-retest or inter-rater reliabilityThe ICC is the share of variance due to real between-subject differences rather than measurement error.
When NOT to — use instead
- Ratings are categoricalFor nominal judgements use a chance-corrected agreement index. → Cohen's kappa
- Ordered categoriesFor a few ordinal grades, weighted kappa is more natural. → Weighted kappa
Assumptions (and what to do if they fail)
Check: One-way vs two-way, single vs average, and agreement vs consistency give different numbers — pick the form that matches the design (e.g. ICC(2,1) for generalisable single raters measuring absolute agreement).
If violated: The wrong form can move the estimate substantially; report the form explicitly (Shrout-Fleiss / McGraw-Wong notation).
Check: ICC is variance-based and assumes interval-level, roughly normal ratings.
If violated: Heavy skew or floor/ceiling effects distort the variance components.
Check: A consistency ICC ignores systematic rater differences; an agreement ICC penalises them — state which you want.
If violated: Two raters who rank identically but score 1 point apart look perfect on consistency and poor on agreement; the wrong choice misleads.
Ready to run a Intraclass Correlation (ICC) on your own data?
Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.
Run this test →