Internal Consistency and Coefficient Alpha

Computing coefficient alpha from item and total-score variances, the model assumptions that decide whether it equals reliability, understates it, or can exceed it, and why the coefficient rises with test length whatever the items measure and reports nothing about how many dimensions they span.

Definition

Coefficient alpha estimates reliability from a single administration, using the k item variances s i 2 and the variance s T 2 of the total score:

α = k k − 1 ( 1 − ∑ i = 1 k s i 2 s T 2 ) .

The reliability parameter alpha is being compared against is ρ X X ′ = σ T 2 / σ X 2 for the total score X , under the classical model X = T + E with E [ E ] = 0 , E uncorrelated with T , and errors uncorrelated across items. That last assumption is doing as much work as the others and is the one most often violated in practice.

Under that model:

  • Essential tau-equivalence means each item's true score differs from any other's by an additive constant, T i = T j + c i j . When it holds, α = ρ X X ′ exactly.
  • Congeneric items, where true scores are linearly but not additively related, make α ≤ ρ X X ′ : alpha understates reliability, and it is in this sense that alpha is a lower bound.
  • Correlated errors break the derivation itself. The uncorrelated-error assumption is what makes ∑ i σ E i 2 the error variance of the total; positively correlated errors inflate σ X 2 through a term the formula attributes to true-score variance, so α can exceed ρ X X ′ . The lower-bound reading is not available here.

So "alpha is a lower bound on reliability" is conditional, not unconditional: it holds under the classical model with uncorrelated errors, and it is then a lower bound because tau-equivalence may fail. A low alpha is ambiguous between unreliable items and a violated tau-equivalence assumption; a high alpha is ambiguous between genuine reliability and correlated errors.

Test length enters directly. With k items of mean inter-item correlation r ¯ , the Spearman-Brown relation gives

α k = k r ¯ 1 + ( k − 1 ) r ¯ ,

which rises toward 1 as k grows for any fixed r ¯ > 0 .

Assumptions and scope

  • Alpha estimates ρ X X ′ under the classical model with errors uncorrelated across items. Given that, alpha equals reliability under essential tau-equivalence and falls below it for congeneric items, which is the sense in which it is a lower bound. If errors are correlated the bound fails in the other direction and alpha can exceed reliability, so the lower-bound reading must not be quoted unconditionally.

  • The Spearman-Brown relation assumes the added items are comparable to the existing ones. Lengthening a test with items of lower quality does not deliver the predicted gain.

Worked material

Example

Four matrices, four coefficients

Each scale below is six items on ten respondents. The coefficients differ for four different reasons, only one of which is item quality.

---

1. Six items, one construct, strong agreement. Mean inter-item correlation r ¯ = 0.62 .

α = 6 ( 0.62 ) 1 + 5 ( 0.62 ) = 3.72 4.10 = 0.9073 .

This is the case the coefficient is built for: items sharing most of their variation, and a value that reflects it.

2. Twenty items, one construct, weak agreement. r ¯ = 0.20 .

α = 20 ( 0.20 ) 1 + 19 ( 0.20 ) = 4.00 4.80 = 0.8333 .

Items sharing a fifth of their variation, reported as good internal consistency. Compare with case 1: the coefficients are within 0.07 of each other and describe instruments that are not comparable. Reporting k alongside the coefficient is what separates them, and inverting Spearman-Brown recovers r ¯ from any reported pair.

3. Two facets, each coherent, weakly related to each other. Items 1–3 correlate 0.55 within, items 4–6 correlate 0.55 within, and across the two blocks the correlation is 0.05 .

The mean over all 15 pairs is

r ¯ = 3 ( 0.55 ) + 3 ( 0.55 ) + 9 ( 0.05 ) 15 = 1.65 + 1.65 + 0.45 15 = 0.25 ,

giving α = 6 ( 0.25 ) 1 + 5 ( 0.25 ) = 0.6667 .

A middling coefficient produced by a genuinely two-dimensional instrument. Deleting items until it rises would raise the statistic by discarding one facet. The coefficient cannot distinguish this case from case 4.

4. Six items, one construct, two of them poorly written. Four items correlate 0.55 with each other; the remaining two correlate about 0.08 with everything.

r ¯ = 6 ( 0.55 ) + 9 ( 0.08 ) 15 = 3.30 + 0.72 15 = 0.268 , α = 0.6873 .

Nearly the same coefficient as case 3, and the right response is the opposite: here the two weak items should be revised or removed, because they measure nothing. In case 3 removing them destroys content.

---

What separates cases 3 and 4 is the pattern, not the summary. In case 3 the weak correlations are structured: low across blocks, high within. In case 4 they are diffuse: two items weak against everything. A single number averages over that distinction, and only the matrix shows it, which is why examining the correlation structure is the step the coefficient cannot replace.

A check worth running on any reported alpha. Invert Spearman-Brown to get the implied r ¯ , then compare it against the matrix. Agreement confirms the arithmetic; disagreement means the reported figure and the reported structure describe different data.

Contrast

Pairs that differ in one respect

The same coefficient from four items and from forty.

4 items40 items
α 0.9575 0.9091
implied r ¯ 0.8492 0.20
items sharefive sixths of their variationa fifth

The two coefficients are close and the instruments are not comparable. Reported as a single number with no k , they are indistinguishable, which is why the number of items belongs beside the coefficient.

The coefficient against the standard error of measurement.

α = 0.9575 sounds like precision. On this scale σ X = 5.2715 , so SEM = 1.0870 and an observed total carries about ± 2.13 points at 95%. On a test scored 0 to 16, respondents scoring 11 and 13 cannot be separated. The coefficient is a unitless ratio; only the second figure answers a question about a person.

Alpha falling because items are bad, against alpha falling because a second dimension arrived.

Adding two items that measure something else drops α from 0.9575 to 0.5097 . Nothing is wrong with those items; their mean correlation with the original four is − 0.3092 , and negative covariances suppress the total variance the coefficient depends on. Deleting items until alpha recovers would optimise the statistic and discard half a two-dimensional construct. The two causes produce the same symptom and call for opposite responses.

Reliability restricted by population, against reliability changed by the instrument.

narrower groupshortened test
σ T 2 fallsfalls
σ E 2 unchangedrises per item removed
what changedwho was testedthe instrument

Both lower the coefficient and only the second is a fact about the test. A figure quoted without its population cannot be assigned to either.

Reliability against validity, on a depression screen.

Consistency of 0.94 with a correlation of 0.12 against clinical diagnosis. The first is a property of the scores in that sample; the second bears on an interpretation. Reliability bounds validity, the correlation cannot exceed what reliability permits, and supplies none of it, so a high coefficient is a precondition rather than evidence.

Common errors

Common misconception

That a high coefficient alpha shows a scale measures a single construct, and that a low one shows it does not. Alpha responds to the number of items as well as to their intercorrelation: by the Spearman-Brown relation, forty items whose average pairwise correlation is only 0.20 give α = 0.9091 , which is reported as excellent for items sharing a fifth of their variation. It moves in the other direction too. Four coherent items with mean inter-item correlation 0.8508 give α = 0.9575 ; adding two items that correlate negatively with them, at a mean cross-correlation of − 0.3092 , drops alpha to 0.5097 . Neither a high value nor a low one settles how many dimensions are present, and establishing that requires examining the correlation structure rather than a single coefficient.

Related units

Requires

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.