Reliability and Measurement Error
An observed score as a true score plus error, reliability as the share of observed variance that is not error, the standard error of measurement that puts the same information on the score scale, and why both are properties of scores from a population rather than of the instrument.
Definition
Classical test theory writes an observed score as
Reliability is the share of observed variance that is true-score variance,
a number between 0 and 1. The standard error of measurement converts a reliability coefficient back to the score scale:
Assumptions and scope
The decomposition
defines the true score as an expectation over hypothetical repeated administrations under identical conditions. It is not the respondent's real standing on the construct, and nothing in classical test theory claims otherwise. requires a reliability coefficient. Substitutinguses alpha as the reliability estimate under the assumed measurement model, which is exact only under essential tau-equivalence with uncorrelated errors; where those fail, the resulting SEM inherits the error in alpha. A
band is a 95% interval only under an assumed error distribution, conventionally normal, and it describes the spread of observed scores around a fixed true score over hypothetical repeated administrations. It is not a 95% interval for the respondent's true score without further assumptions, and reporting it as one overstates what the measurement model supplies.Reliability coefficients are population-dependent. Restricting the range of the group being tested lowers the coefficient without changing the instrument, which is why a figure quoted from a manual cannot be assumed to hold for a new group.
Worked material
Example
Five instruments and what each coefficient settles
A 40-item anxiety questionnaire reporting
A 4-item scale reporting
A classroom test where
---
In all three the coefficient is doing its job and is being read without the context that makes it interpretable: the number of items, and the population.