Intraclass Correlation (ICC)

Verified

Reliability (legacy hub)

Independently verified. Every statistic this test reports has been re-derived against an independent reference — never the library the pipeline itself calls — the rendered output was read back in a browser, and the result is locked with a committed regression suite.

Run this test straight away on a free built-in teaching dataset — no data of your own needed — or bring your own. Either opens the guided workspace: variable setup, assumption diagnostics, results with effect sizes and confidence intervals, figures, and APA-ready reporting.

Loading teaching datasets…

Or use your own dataset

Loading your datasets…

Intraclass correlation — agreement or consistency among repeated measurements (multiple raters, or the same measure over time).

The ICC is the share of total variance due to real differences between subjects rather than measurement error. Different forms (ICC(1), (2,1), (3,k)) suit rater vs test-retest designs; ≥ .75 is good.

Worked example

Do three clinicians agree when rating the same patients?

Three raters scored 40 patients on a continuous scale; a two-way ICC quantifies their agreement.

Result

Agreement was good, ICC(2,1) = .76, 95% CI [.64, .86].

How you'd report it (APA)

Inter-rater agreement was good, ICC(2,1) = .76, 95% CI [.64, .86].

Try it yourself: Load this ready-made sample and follow the run above.

When to use it

  • Agreement on continuous measurements
    Several raters score the same subjects on a numeric scale, or the same instrument is repeated over time. e.g. 3 clinicians rating 40 patients.
  • Test-retest or inter-rater reliability
    The ICC is the share of variance due to real between-subject differences rather than measurement error.

When NOT to — use instead

  • Ratings are categorical
    For nominal judgements use a chance-corrected agreement index. Cohen's kappa
  • Ordered categories
    For a few ordinal grades, weighted kappa is more natural. Weighted kappa

Assumptions (and what to do if they fail)

Choose the correct ICC formhigh

Check: One-way vs two-way, single vs average, and agreement vs consistency give different numbers — pick the form that matches the design (e.g. ICC(2,1) for generalisable single raters measuring absolute agreement).

If violated: The wrong form can move the estimate substantially; report the form explicitly (Shrout-Fleiss / McGraw-Wong notation).

Approximately normal, continuous scoresmedium

Check: ICC is variance-based and assumes interval-level, roughly normal ratings.

If violated: Heavy skew or floor/ceiling effects distort the variance components.

Consistency vs absolute agreementmedium

Check: A consistency ICC ignores systematic rater differences; an agreement ICC penalises them — state which you want.

If violated: Two raters who rank identically but score 1 point apart look perfect on consistency and poor on agreement; the wrong choice misleads.

Ready to run a Intraclass Correlation (ICC) on your own data?

Guided setup, automatic assumption checks, effect sizes, figures and an APA write-up.

Run this test →