Internal Consistency and Coefficient Alpha

What you will be able to do

Given item and total-score variances, the learner can compute coefficient alpha, state the measurement model under which it estimates reliability and when the lower-bound reading is available, apply the Spearman-Brown relation to separate item quality from test length, and decline to read dimensionality off the coefficient.

Orientation

A coefficient that rises with length

The number belonging to the first question is the one most often misread. Coefficient alpha is computed from item variances and the variance of the total, and it responds to how many items there are as well as to how well they hang together. Forty items whose average pairwise correlation is only 0.20 reach α = 0.9091 . A figure usually reported as excellent for a set of items sharing a fifth of their variation. Raising alpha by lengthening a test is always available and improves nothing.

This unit covers computing the coefficient, the model under which it estimates reliability, and why it reports nothing about how many dimensions the items span.

Definition

What alpha estimates, and under which model

What alpha estimates, and under which model. The target is ρ X X ′ = σ T 2 / σ X 2 for the total score, under the classical model with errors uncorrelated across items. Given that model, the relation between α and ρ X X ′ depends on how the items relate to one another:

Item structureRelationReading
essentially tau-equivalent α = ρ X X ′ exact
congeneric α ≤ ρ X X ′ alpha understates: the lower-bound case
errors correlatedbound fails α can exceed ρ X X ′

So "alpha is a lower bound" is a statement within the uncorrelated-error model, where it is a bound because tau-equivalence may fail. Positively correlated errors, from adjacent similarly-worded items or a shared response set, inflate σ X 2 through a term the formula credits to true-score variance, and alpha then overstates reliability.

The practical consequence is that both directions are ambiguous. A low alpha does not separate genuinely inconsistent items from a violated tau-equivalence assumption. A high alpha does not separate genuine reliability from correlated errors, which is why the item wording and administration conditions bear on whether the figure can be read as a bound at all.

Intuition

Why the coefficient can be raised without improving anything

Alpha compares the sum of the item variances against the variance of the total score. If the items are unrelated, the total's variance is roughly the sum of the parts and the ratio sits near 1, so alpha is near 0. If the items move together, the total's variance is inflated by their covariances, the ratio falls, and alpha rises. The coefficient is therefore reading how much of the total variance comes from items agreeing with each other.

That quantity has two inputs, and only one of them is item quality. Covariance terms grow as k 2 while item variances grow as k , so lengthening a test raises alpha for any fixed level of agreement. The Spearman-Brown relation makes it exact:

k α at r ¯ = 0.20
4 0.5000
10 0.7143
20 0.8333
40 0.9091

Every row describes items sharing a fifth of their variation. The last row would be reported as excellent internal consistency. Nothing about the items changed between rows.

And the coefficient can move the other way for reasons that are not item quality either. Four coherent items with mean inter-item correlation 0.8508 give α = 0.9575 . Add two items that correlate negatively with them, mean cross-correlation − 0.3092 , and alpha falls to 0.5097 . The two added items are not defective; they measure something else, and their negative covariances with the first four suppress the total variance that alpha depends on.

So a high value does not establish one dimension and a low value does not establish several. The coefficient is not answering that question at all, and answering it means looking at the correlation structure.

Example

Four matrices, four coefficients

Each scale below is six items on ten respondents. The coefficients differ for four different reasons, only one of which is item quality.

---

1. Six items, one construct, strong agreement. Mean inter-item correlation r ¯ = 0.62 .

α = 6 ( 0.62 ) 1 + 5 ( 0.62 ) = 3.72 4.10 = 0.9073 .

This is the case the coefficient is built for: items sharing most of their variation, and a value that reflects it.

2. Twenty items, one construct, weak agreement. r ¯ = 0.20 .

α = 20 ( 0.20 ) 1 + 19 ( 0.20 ) = 4.00 4.80 = 0.8333 .

Items sharing a fifth of their variation, reported as good internal consistency. Compare with case 1: the coefficients are within 0.07 of each other and describe instruments that are not comparable. Reporting k alongside the coefficient is what separates them, and inverting Spearman-Brown recovers r ¯ from any reported pair.

3. Two facets, each coherent, weakly related to each other. Items 1–3 correlate 0.55 within, items 4–6 correlate 0.55 within, and across the two blocks the correlation is 0.05 .

The mean over all 15 pairs is

r ¯ = 3 ( 0.55 ) + 3 ( 0.55 ) + 9 ( 0.05 ) 15 = 1.65 + 1.65 + 0.45 15 = 0.25 ,

giving α = 6 ( 0.25 ) 1 + 5 ( 0.25 ) = 0.6667 .

A middling coefficient produced by a genuinely two-dimensional instrument. Deleting items until it rises would raise the statistic by discarding one facet. The coefficient cannot distinguish this case from case 4.

4. Six items, one construct, two of them poorly written. Four items correlate 0.55 with each other; the remaining two correlate about 0.08 with everything.

r ¯ = 6 ( 0.55 ) + 9 ( 0.08 ) 15 = 3.30 + 0.72 15 = 0.268 , α = 0.6873 .

Nearly the same coefficient as case 3, and the right response is the opposite: here the two weak items should be revised or removed, because they measure nothing. In case 3 removing them destroys content.

---

What separates cases 3 and 4 is the pattern, not the summary. In case 3 the weak correlations are structured: low across blocks, high within. In case 4 they are diffuse: two items weak against everything. A single number averages over that distinction, and only the matrix shows it, which is why examining the correlation structure is the step the coefficient cannot replace.

A check worth running on any reported alpha. Invert Spearman-Brown to get the implied r ¯ , then compare it against the matrix. Agreement confirms the arithmetic; disagreement means the reported figure and the reported structure describe different data.

Procedure

Computing the coefficient and reporting it usefully

To compute coefficient alpha.

  1. Score every item in the same direction. Reverse-keyed items must be reversed first; forgetting this produces negative covariances and a coefficient that is meaninglessly low.
  2. Compute each item's variance across respondents, using one denominator consistently.
  3. Compute each respondent's total and the variance of those totals.
  4. Apply α = k k − 1 ( 1 − ∑ i s i 2 / s T 2 ) .
  5. Check the result against the structure. Inverting Spearman-Brown gives the mean inter-item correlation the coefficient implies; comparing that against the correlations actually observed catches arithmetic errors and reveals whether the value came from strong items or from many of them.

To report it usefully.

  1. Convert to the score scale: SEM = σ X 1 − α .
  2. State the band an observed score carries, roughly ± 2 SEM for 95%.
  3. Name the population and conditions the coefficient was computed under. A figure without them describes nothing transferable.
  4. Report k alongside the coefficient, since the same value means different things at four items and at forty.

To examine dimensionality, which alpha does not settle. Inspect the inter-item correlation matrix for blocks, or run a factor analysis. A single dominant factor supports a total score; two blocks argue for two subscale scores, and summing across them produces a number whose meaning is unclear whatever alpha says.

Worked example

One matrix, four readings

Ten respondents answer six items scored 0 to 4. Items 1 to 4 were written for one construct; items 5 and 6 were written for a different one.

RespondentI1I2I3I4I5I6
1443410
2333243
3212204
4434321
5110134
6223240
7343402
8011023
9444331
10101112

Step 1: alpha on items 1 to 4.

Item variances (divisor n − 1 = 9 ): 2.0444 , 2.2333 , 1.8222 , 1.7333 , summing to 7.8333 . The four-item totals have variance s T 2 = 2501 / 90 = 27.7889 .

α = 4 3 ( 1 − 7.8333 27.7889 ) = 7184 7503 = 0.9575 .

A consistency check worth doing. Inverting Spearman-Brown, α = 0.9575 at k = 4 implies a mean inter-item correlation of r ¯ = 0.8492 . Computing the six pairwise correlations directly gives 0.8736 , 0.8865 , 0.8972 , 0.8152 , 0.8697 , 0.7627 , whose mean is 0.8508 . The two agree to three decimals, so the formula and the matrix are describing the same structure.

Step 2: the coefficient on the score scale. σ X = 27.7889 = 5.2715 , so

SEM = 5.2715 1 − 0.9575 = 1.0870 .

On a total running 0 to 16, an observed score carries a 95% band of roughly ± 2.13 points. A respondent scoring 11 and one scoring 13 are not distinguishable by this instrument. The coefficient of 0.9575 conveys none of that.

Step 3: add the two items from the other construct.

α six items = 0.5097 , a fall of  0.4477 .

The mean correlation between the two groups is − 0.3092 , ranging from − 0.6626 to + 0.1104 , and r ( I5 , I6 ) = − 0.1500 . Items 5 and 6 are not defective items; they measure something else, and their negative covariances with the first four suppress the total variance alpha is computed from.

The lesson is not "remove items that lower alpha". That rule would optimise the coefficient rather than the measurement, and on a genuinely two-dimensional construct it discards half the content. The finding here is that alpha moved a great deal because the structure changed, which is a reason to examine the structure.

Step 4: the same coefficient reached by length instead. Suppose forty items with mean inter-item correlation r ¯ = 0.20 :

α 40 = 40 ( 0.20 ) 1 + 39 ( 0.20 ) = 0.9091 .

This is close to the 0.9575 of step 1 and describes a completely different instrument, items sharing a fifth of their variation rather than five sixths. Reported as a single number, the two are indistinguishable.

---

What none of this establishes. Every figure above concerns consistency. Whether these scores should inform any particular decision is untouched by all four steps, and settling it requires stating what the scores are to mean and gathering evidence on that claim.

Contrast

Pairs that differ in one respect

The same coefficient from four items and from forty.

4 items40 items
α 0.9575 0.9091
implied r ¯ 0.8492 0.20
items sharefive sixths of their variationa fifth

The two coefficients are close and the instruments are not comparable. Reported as a single number with no k , they are indistinguishable, which is why the number of items belongs beside the coefficient.

The coefficient against the standard error of measurement.

α = 0.9575 sounds like precision. On this scale σ X = 5.2715 , so SEM = 1.0870 and an observed total carries about ± 2.13 points at 95%. On a test scored 0 to 16, respondents scoring 11 and 13 cannot be separated. The coefficient is a unitless ratio; only the second figure answers a question about a person.

Alpha falling because items are bad, against alpha falling because a second dimension arrived.

Adding two items that measure something else drops α from 0.9575 to 0.5097 . Nothing is wrong with those items; their mean correlation with the original four is − 0.3092 , and negative covariances suppress the total variance the coefficient depends on. Deleting items until alpha recovers would optimise the statistic and discard half a two-dimensional construct. The two causes produce the same symptom and call for opposite responses.

Reliability restricted by population, against reliability changed by the instrument.

narrower groupshortened test
σ T 2 fallsfalls
σ E 2 unchangedrises per item removed
what changedwho was testedthe instrument

Both lower the coefficient and only the second is a fact about the test. A figure quoted without its population cannot be assigned to either.

Reliability against validity, on a depression screen.

Consistency of 0.94 with a correlation of 0.12 against clinical diagnosis. The first is a property of the scores in that sample; the second bears on an interpretation. Reliability bounds validity, the correlation cannot exceed what reliability permits, and supplies none of it, so a high coefficient is a precondition rather than evidence.

Warning

Optimising the coefficient rather than the measurement

Lengthening the test to reach a threshold. Alpha rises with k at any fixed level of item agreement. Moving from 10 items to 40 at r ¯ = 0.20 takes the coefficient from 0.7143 to 0.9091 without a single item improving. Where a publication norm sets a threshold at 0.80 or 0.90 , adding near-duplicate items clears it, and the instrument is longer rather than better.

Deleting items until alpha recovers. Alpha fell from 0.9575 to 0.5097 when two items measuring a second construct were added. Removing them restores the coefficient and removes the second dimension from the instrument. If the construct genuinely has two facets, the procedure has optimised a statistic by discarding content, and nothing in the coefficient signals that this is what happened.

Reading alpha as evidence of one dimension. It is not, in either direction. A long test of weakly related items scores high, and a genuinely two-dimensional instrument can score low. Establishing dimensionality means inspecting the correlation structure.

Next step

Practice Internal Consistency and Coefficient Alpha

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.