Reliability and Measurement Error
What you will be able to do
Given a reliability coefficient and the score standard deviation, the learner can compute the standard error of measurement, state the band an individual score carries and the assumptions that band rests on, and attribute the coefficient to scores from a population rather than to the instrument.
Orientation
A stable answer, and what stability is a property of
Any measurement invites two questions. Does the instrument give a stable answer? And does that answer mean what it is being taken to mean?
They are independent. A bathroom scale that reads three kilograms heavy gives the same answer every time and is stable and wrong. A test can produce highly consistent scores that predict nothing about the decision they are being used to inform.
The first question is reliability, and it has a number attached. The second is validity, and it does not: it is an argument about a particular interpretation of scores for a particular use, and the same instrument can be validly interpreted one way and not another.
This unit takes the first question. Reliability is a ratio of variances, both computed on a group, so it describes an administration rather than a product. The standard error of measurement puts the same information on the score scale, which is what a decision about one person actually needs.
Definition
What each quantity is defined over
The canonical statements above give the decomposition, the coefficient and the relation to test length. What follows is the scope of each, which is where the misreadings start.
The true score is an expectation, not a fact about the person.
Reliability is a ratio of two variances, both computed on a group.
| Quantity | Defined over | Changes when |
|---|---|---|
| the tested population | the group's spread changes | |
| administrations | conditions or the instrument change | |
| both | either changes |
This is why a coefficient does not transfer between groups. Restrict the range, test only the top quartile, and
The standard error of measurement is the same information on the score scale.
Two qualifications travel with it. Substituting
Intuition
Why reliability belongs to scores
Why reliability belongs to scores rather than to instruments. Reliability is the share of observed variance that is true-score variance. Test a narrower group and the true-score variance falls while the error variance does not, so the coefficient drops although the instrument is unchanged. A figure quoted from a manual describes the manual's sample.
Why the standard error of measurement is the more useful report. For the four coherent items,
Example
Five instruments and what each coefficient settles
A 40-item anxiety questionnaire reporting
A 4-item scale reporting
A classroom test where
---
In all three the coefficient is doing its job and is being read without the context that makes it interpretable: the number of items, and the population.
Warning
Overreading a reliability coefficient
Quoting a coefficient from a manual. Reliability is a ratio of variances computed on a group. Administering the same instrument to a narrower group lowers it, since true-score variance falls while error variance does not. A figure without its population and conditions describes an administration that is not the one being conducted.
Reporting the coefficient and not the standard error of measurement.