Internal Consistency and Coefficient Alpha
Computing coefficient alpha from item and total-score variances, the model assumptions that decide whether it equals reliability, understates it, or can exceed it, and why the coefficient rises with test length whatever the items measure and reports nothing about how many dimensions they span.
Definition
Coefficient alpha estimates reliability from a single administration, using the
The reliability parameter alpha is being compared against is
Under that model:
- Essential tau-equivalence means each item's true score differs from any other's by an additive constant,
. When it holds, exactly. - Congeneric items, where true scores are linearly but not additively related, make
: alpha understates reliability, and it is in this sense that alpha is a lower bound. - Correlated errors break the derivation itself. The uncorrelated-error assumption is what makes
the error variance of the total; positively correlated errors inflate through a term the formula attributes to true-score variance, so can exceed . The lower-bound reading is not available here.
So "alpha is a lower bound on reliability" is conditional, not unconditional: it holds under the classical model with uncorrelated errors, and it is then a lower bound because tau-equivalence may fail. A low alpha is ambiguous between unreliable items and a violated tau-equivalence assumption; a high alpha is ambiguous between genuine reliability and correlated errors.
Test length enters directly. With
which rises toward 1 as
Assumptions and scope
Alpha estimates
under the classical model with errors uncorrelated across items. Given that, alpha equals reliability under essential tau-equivalence and falls below it for congeneric items, which is the sense in which it is a lower bound. If errors are correlated the bound fails in the other direction and alpha can exceed reliability, so the lower-bound reading must not be quoted unconditionally. The Spearman-Brown relation assumes the added items are comparable to the existing ones. Lengthening a test with items of lower quality does not deliver the predicted gain.
Worked material
Example
Four matrices, four coefficients
Each scale below is six items on ten respondents. The coefficients differ for four different reasons, only one of which is item quality.
---
1. Six items, one construct, strong agreement. Mean inter-item correlation
This is the case the coefficient is built for: items sharing most of their variation, and a value that reflects it.
2. Twenty items, one construct, weak agreement.
Items sharing a fifth of their variation, reported as good internal consistency. Compare with case 1: the coefficients are within
3. Two facets, each coherent, weakly related to each other. Items 1–3 correlate
The mean over all 15 pairs is
giving
A middling coefficient produced by a genuinely two-dimensional instrument. Deleting items until it rises would raise the statistic by discarding one facet. The coefficient cannot distinguish this case from case 4.
4. Six items, one construct, two of them poorly written. Four items correlate
Nearly the same coefficient as case 3, and the right response is the opposite: here the two weak items should be revised or removed, because they measure nothing. In case 3 removing them destroys content.
---
What separates cases 3 and 4 is the pattern, not the summary. In case 3 the weak correlations are structured: low across blocks, high within. In case 4 they are diffuse: two items weak against everything. A single number averages over that distinction, and only the matrix shows it, which is why examining the correlation structure is the step the coefficient cannot replace.
A check worth running on any reported alpha. Invert Spearman-Brown to get the implied
Contrast
Pairs that differ in one respect
The same coefficient from four items and from forty.
| 4 items | 40 items | |
|---|---|---|
| implied | ||
| items share | five sixths of their variation | a fifth |
The two coefficients are close and the instruments are not comparable. Reported as a single number with no
The coefficient against the standard error of measurement.
Alpha falling because items are bad, against alpha falling because a second dimension arrived.
Adding two items that measure something else drops
Reliability restricted by population, against reliability changed by the instrument.
| narrower group | shortened test | |
|---|---|---|
| falls | falls | |
| unchanged | rises per item removed | |
| what changed | who was tested | the instrument |
Both lower the coefficient and only the second is a fact about the test. A figure quoted without its population cannot be assigned to either.
Reliability against validity, on a depression screen.
Consistency of
Common errors
Common misconception
That a high coefficient alpha shows a scale measures a single construct, and that a low one shows it does not. Alpha responds to the number of items as well as to their intercorrelation: by the Spearman-Brown relation, forty items whose average pairwise correlation is only
Related units
Requires
Learn this topic
Used in
Sources
- Coefficient Alpha and the Internal Structure of Tests (1951)
- Statistical Theories of Mental Test Scores (1968)