Module 1 of 1 · Lesson 1 of 3

Reliability and Measurement Error

Reliability as a ratio of variances, and the standard error of measurement on the score scale.

What you will be able to do

Given a reliability coefficient and the score standard deviation, the learner can compute the standard error of measurement, state the band an individual score carries and the assumptions that band rests on, and attribute the coefficient to scores from a population rather than to the instrument.

Orientation

A stable answer, and what stability is a property of

Any measurement invites two questions. Does the instrument give a stable answer? And does that answer mean what it is being taken to mean?

They are independent. A bathroom scale that reads three kilograms heavy gives the same answer every time and is stable and wrong. A test can produce highly consistent scores that predict nothing about the decision they are being used to inform.

The first question is reliability, and it has a number attached. The second is validity, and it does not: it is an argument about a particular interpretation of scores for a particular use, and the same instrument can be validly interpreted one way and not another.

This unit takes the first question. Reliability is a ratio of variances, both computed on a group, so it describes an administration rather than a product. The standard error of measurement puts the same information on the score scale, which is what a decision about one person actually needs.

Definition

What each quantity is defined over

The canonical statements above give the decomposition, the coefficient and the relation to test length. What follows is the scope of each, which is where the misreadings start.

The true score is an expectation, not a fact about the person. T is defined as the mean observed score over hypothetical independent administrations under identical conditions. It is a property of the person and the instrument together. Nothing in classical test theory claims T is the respondent's real standing on the construct, if the instrument measures the wrong thing, T is the stable value of the wrong measurement.

Reliability is a ratio of two variances, both computed on a group.

QuantityDefined overChanges when
σ T 2 the tested populationthe group's spread changes
σ E 2 administrationsconditions or the instrument change
ρ X X ′ botheither changes

This is why a coefficient does not transfer between groups. Restrict the range, test only the top quartile, and σ T 2 falls while σ E 2 does not, so the ratio drops. The instrument is identical; the reported reliability is lower, and correctly so.

The standard error of measurement is the same information on the score scale. SEM = σ X 1 − ρ X X ′ answers what a reader deciding about an individual needs: how far an observed score might sit from the true one. A coefficient is a unitless ratio and cannot answer that.

Two qualifications travel with it. Substituting α for ρ X X ′ uses alpha as the reliability estimate under the assumed measurement model, so the SEM inherits whatever error that substitution carries. And a ± 2 SEM band is a 95% interval only under an assumed error distribution, conventionally normal; it describes the spread of observed scores about a fixed true score over hypothetical repeated administrations, which is not the same as a 95% interval for the respondent's true score.

Intuition

Why reliability belongs to scores

Why reliability belongs to scores rather than to instruments. Reliability is the share of observed variance that is true-score variance. Test a narrower group and the true-score variance falls while the error variance does not, so the coefficient drops although the instrument is unchanged. A figure quoted from a manual describes the manual's sample.

Why the standard error of measurement is the more useful report. For the four coherent items, α = 0.9575 and σ X = 5.2715 give SEM = 1.0870 on a total running 0 to 16. An individual's observed score is worth about ± 2.1 points at 95% confidence. The coefficient sounds like precision; the ± 2.1 is the precision, and it is what a decision about a person has to be made against.

Example

Five instruments and what each coefficient settles

A 40-item anxiety questionnaire reporting α = 0.91 . Consistent with items sharing a great deal, and equally consistent with items sharing a fifth of their variation: at r ¯ = 0.20 , forty items give 0.9091 . The coefficient alone does not distinguish the two, and reporting k alongside it does most of the work of telling them apart.

A 4-item scale reporting α = 0.96 . Far more informative, because four items cannot reach that by length. Inverting Spearman-Brown, the implied mean inter-item correlation is about 0.85 . The items are nearly redundant with one another, which raises a different question: whether four items this similar are covering the construct or asking one question four ways.

A classroom test where α drops from 0.82 to 0.61 between the full cohort and the top set. Nothing about the test changed. Restricting the range cut true-score variance while error variance stayed put, so the ratio fell. The second figure is the correct reliability for that group, and the instrument is neither better nor worse than it was.

---

In all three the coefficient is doing its job and is being read without the context that makes it interpretable: the number of items, and the population.

Warning

Overreading a reliability coefficient

Quoting a coefficient from a manual. Reliability is a ratio of variances computed on a group. Administering the same instrument to a narrower group lowers it, since true-score variance falls while error variance does not. A figure without its population and conditions describes an administration that is not the one being conducted.

Reporting the coefficient and not the standard error of measurement. α = 0.9575 conveys precision; the corresponding SEM = 1.0870 says an individual's total is worth ± 2.13 points at 95% on a 0 to 16 scale. Decisions about people are made on the second figure, and it is the one usually omitted.

Next step

Practice Reliability and Measurement Error

Practice records what support you used, so the evidence reflects how you actually performed.

Practice this lessonSkip to Internal Consistency and Coefficient Alpha

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.