Module 1 of 1 · Lesson 2 of 3
Internal Consistency and Coefficient Alpha
Coefficient alpha, the model it assumes, and why it rises with test length.
What you will be able to do
Given item and total-score variances, the learner can compute coefficient alpha, state the measurement model under which it estimates reliability and when the lower-bound reading is available, apply the Spearman-Brown relation to separate item quality from test length, and decline to read dimensionality off the coefficient.
Orientation
A coefficient that rises with length
The number belonging to the first question is the one most often misread. Coefficient alpha is computed from item variances and the variance of the total, and it responds to how many items there are as well as to how well they hang together. Forty items whose average pairwise correlation is only
This unit covers computing the coefficient, the model under which it estimates reliability, and why it reports nothing about how many dimensions the items span.
Definition
What alpha estimates, and under which model
What alpha estimates, and under which model. The target is
| Item structure | Relation | Reading |
|---|---|---|
| essentially tau-equivalent | exact | |
| congeneric | alpha understates: the lower-bound case | |
| errors correlated | bound fails |
So "alpha is a lower bound" is a statement within the uncorrelated-error model, where it is a bound because tau-equivalence may fail. Positively correlated errors, from adjacent similarly-worded items or a shared response set, inflate
The practical consequence is that both directions are ambiguous. A low alpha does not separate genuinely inconsistent items from a violated tau-equivalence assumption. A high alpha does not separate genuine reliability from correlated errors, which is why the item wording and administration conditions bear on whether the figure can be read as a bound at all.
Intuition
Why the coefficient can be raised without improving anything
Alpha compares the sum of the item variances against the variance of the total score. If the items are unrelated, the total's variance is roughly the sum of the parts and the ratio sits near 1, so alpha is near 0. If the items move together, the total's variance is inflated by their covariances, the ratio falls, and alpha rises. The coefficient is therefore reading how much of the total variance comes from items agreeing with each other.
That quantity has two inputs, and only one of them is item quality. Covariance terms grow as
| 4 | |
| 10 | |
| 20 | |
| 40 |
Every row describes items sharing a fifth of their variation. The last row would be reported as excellent internal consistency. Nothing about the items changed between rows.
And the coefficient can move the other way for reasons that are not item quality either. Four coherent items with mean inter-item correlation
So a high value does not establish one dimension and a low value does not establish several. The coefficient is not answering that question at all, and answering it means looking at the correlation structure.
Example
Four matrices, four coefficients
Each scale below is six items on ten respondents. The coefficients differ for four different reasons, only one of which is item quality.
---
1. Six items, one construct, strong agreement. Mean inter-item correlation
This is the case the coefficient is built for: items sharing most of their variation, and a value that reflects it.
2. Twenty items, one construct, weak agreement.
Items sharing a fifth of their variation, reported as good internal consistency. Compare with case 1: the coefficients are within
3. Two facets, each coherent, weakly related to each other. Items 1–3 correlate
The mean over all 15 pairs is
giving
A middling coefficient produced by a genuinely two-dimensional instrument. Deleting items until it rises would raise the statistic by discarding one facet. The coefficient cannot distinguish this case from case 4.
4. Six items, one construct, two of them poorly written. Four items correlate
Nearly the same coefficient as case 3, and the right response is the opposite: here the two weak items should be revised or removed, because they measure nothing. In case 3 removing them destroys content.
---
What separates cases 3 and 4 is the pattern, not the summary. In case 3 the weak correlations are structured: low across blocks, high within. In case 4 they are diffuse: two items weak against everything. A single number averages over that distinction, and only the matrix shows it, which is why examining the correlation structure is the step the coefficient cannot replace.
A check worth running on any reported alpha. Invert Spearman-Brown to get the implied
Procedure
Computing the coefficient and reporting it usefully
To compute coefficient alpha.
- Score every item in the same direction. Reverse-keyed items must be reversed first; forgetting this produces negative covariances and a coefficient that is meaninglessly low.
- Compute each item's variance across respondents, using one denominator consistently.
- Compute each respondent's total and the variance of those totals.
- Apply
. - Check the result against the structure. Inverting Spearman-Brown gives the mean inter-item correlation the coefficient implies; comparing that against the correlations actually observed catches arithmetic errors and reveals whether the value came from strong items or from many of them.
To report it usefully.
- Convert to the score scale:
. - State the band an observed score carries, roughly
for 95%. - Name the population and conditions the coefficient was computed under. A figure without them describes nothing transferable.
- Report
alongside the coefficient, since the same value means different things at four items and at forty.
To examine dimensionality, which alpha does not settle. Inspect the inter-item correlation matrix for blocks, or run a factor analysis. A single dominant factor supports a total score; two blocks argue for two subscale scores, and summing across them produces a number whose meaning is unclear whatever alpha says.
Worked example
One matrix, four readings
Ten respondents answer six items scored 0 to 4. Items 1 to 4 were written for one construct; items 5 and 6 were written for a different one.
| Respondent | I1 | I2 | I3 | I4 | I5 | I6 |
|---|---|---|---|---|---|---|
| 1 | 4 | 4 | 3 | 4 | 1 | 0 |
| 2 | 3 | 3 | 3 | 2 | 4 | 3 |
| 3 | 2 | 1 | 2 | 2 | 0 | 4 |
| 4 | 4 | 3 | 4 | 3 | 2 | 1 |
| 5 | 1 | 1 | 0 | 1 | 3 | 4 |
| 6 | 2 | 2 | 3 | 2 | 4 | 0 |
| 7 | 3 | 4 | 3 | 4 | 0 | 2 |
| 8 | 0 | 1 | 1 | 0 | 2 | 3 |
| 9 | 4 | 4 | 4 | 3 | 3 | 1 |
| 10 | 1 | 0 | 1 | 1 | 1 | 2 |
Step 1: alpha on items 1 to 4.
Item variances (divisor
A consistency check worth doing. Inverting Spearman-Brown,
Step 2: the coefficient on the score scale.
On a total running 0 to 16, an observed score carries a 95% band of roughly
Step 3: add the two items from the other construct.
The mean correlation between the two groups is
The lesson is not "remove items that lower alpha". That rule would optimise the coefficient rather than the measurement, and on a genuinely two-dimensional construct it discards half the content. The finding here is that alpha moved a great deal because the structure changed, which is a reason to examine the structure.
Step 4: the same coefficient reached by length instead. Suppose forty items with mean inter-item correlation
This is close to the
---
What none of this establishes. Every figure above concerns consistency. Whether these scores should inform any particular decision is untouched by all four steps, and settling it requires stating what the scores are to mean and gathering evidence on that claim.
Contrast
Pairs that differ in one respect
The same coefficient from four items and from forty.
| 4 items | 40 items | |
|---|---|---|
| implied | ||
| items share | five sixths of their variation | a fifth |
The two coefficients are close and the instruments are not comparable. Reported as a single number with no
The coefficient against the standard error of measurement.
Alpha falling because items are bad, against alpha falling because a second dimension arrived.
Adding two items that measure something else drops
Reliability restricted by population, against reliability changed by the instrument.
| narrower group | shortened test | |
|---|---|---|
| falls | falls | |
| unchanged | rises per item removed | |
| what changed | who was tested | the instrument |
Both lower the coefficient and only the second is a fact about the test. A figure quoted without its population cannot be assigned to either.
Reliability against validity, on a depression screen.
Consistency of
Warning
Optimising the coefficient rather than the measurement
Lengthening the test to reach a threshold. Alpha rises with
Deleting items until alpha recovers. Alpha fell from
Reading alpha as evidence of one dimension. It is not, in either direction. A long test of weakly related items scores high, and a genuinely two-dimensional instrument can score low. Establishing dimensionality means inspecting the correlation structure.