Validity as an Argument for an Interpretation
What you will be able to do
Given a proposed use of test scores, the learner can state the interpretation the use relies on, name the kinds of evidence that would bear on it, and explain why reliability bounds a validity coefficient without supplying validity and why evidence for one use does not transfer to another.
Orientation
A claim about an interpretation, not about an instrument
Reliability has a number attached. Validity does not, and that difference is structural rather than a gap in the literature.
Validity is the degree to which evidence supports a specific interpretation of scores for a specific use. So the same instrument can be validly interpreted one way and not another, and the phrase "a validated instrument" names no property until the interpretation is stated.
Two consequences follow, and both are routinely violated in practice.
Evidence does not transfer across uses. A reading test validated for identifying pupils who need support carries no validity evidence for ranking schools. The scores are unchanged; the interpretation changed, and the original argument never addressed the new one.
Reliability is necessary and nowhere near sufficient. A score cannot correlate with a criterion more strongly than its own reliability permits, so consistency bounds validity. It supplies none of it: a depression screen with
This unit covers what a validity argument consists of, what kinds of evidence bear on it, and why the argument has to name the decision before it cites a number.
Definition
What a validity claim is defined over
Validity is defined over an interpretation and a use, not over the instrument. The standard formulation asks how far evidence supports a specified interpretation of scores for a specified purpose. Two consequences follow. A test carries no validity in the abstract, so "a validated instrument" names nothing until the interpretation is stated. And evidence gathered for one use does not transfer to another, however similar the scores look.
Reliability bounds validity without supplying it. A score cannot correlate with an external criterion more strongly than its own reliability permits, which makes reliability necessary. It is not sufficient: a perfectly reliable score may be reliably measuring something irrelevant to the decision at hand.
Intuition
What settles a validity claim
Validity is a different kind of claim. Nothing above bears on whether the scores mean what they are being used to mean. That question is settled by stating the interpretation, identifying what would have to hold for it to be sound, and gathering evidence on each, including evidence that could count against it.
Example
One instrument, four proposed uses
A 25-item numeracy assessment,
---
Use 1: telling a teacher which pupils need extra practice. The interpretation is that a low score indicates a pupil who would benefit from additional instruction now. Evidence bearing on it: whether low scorers do improve more than high scorers when given the instruction, and whether the items cover what the instruction addresses. The
Use 2: certifying that a pupil has met a standard. The interpretation is that a score above a cut represents mastery of a defined content domain. Now content evidence becomes central: does the item set sample the domain the standard names, in the proportions the standard implies? A correlation with later attainment is close to irrelevant here, and the operative statistic is classification consistency near the cut, computed from the standard error of measurement rather than from the coefficient.
Use 3: comparing this year's cohort against last year's. The interpretation is that a difference in means reflects a difference in numeracy rather than in the measurement. This requires evidence the two administrations were equivalent: the same items, or equated forms, and no change in who sits the assessment. An instrument perfectly valid for Use 1 supports this only if that equivalence holds, and nothing in the reliability coefficient speaks to it.
Use 4: evaluating teachers by their class's mean score. The interpretation is that class mean reflects teaching quality. This requires everything above plus an argument that intake composition has been accounted for, since a class mean encodes who is in the class as much as what happened in it. Consequential evidence also becomes central: once scores determine evaluations, teaching narrows toward the tested items, which changes what the score measures.
---
| Use | Evidence that matters most | What |
|---|---|---|
| Identify pupils to support | response to instruction | necessary precondition only |
| Certify a standard | content coverage, classification consistency | necessary precondition only |
| Compare cohorts | administration equivalence | necessary precondition only |
| Evaluate teachers | intake adjustment, consequences | necessary precondition only |
The column on the right is the point. A single reliability coefficient contributes the same thing to all four, and it is never the thing in dispute. What differs down the left column is what would have to be shown, and each row is a separate argument that has to be made rather than inherited from the row above.
And the direction of failure. Uses 1 and 2 are the ones the instrument was built for. Use 4 is the furthest from it, and it is the one most likely to be adopted without a new argument, because the scores are already there and already look authoritative.
Procedure
Assembling a validity argument
To assemble a validity argument.
- State the interpretation and the use in one sentence, before any evidence: what the scores are taken to mean, and what decision they will inform.
- List what must hold for that interpretation to be sound, relations with other measures, the content the items cover, the response processes respondents actually use, the consequences of the decision.
- Gather evidence on each, including evidence that could count against the interpretation.
- State the boundary. Evidence supports this interpretation for this use; a different use needs its own argument.
Checks. Confirm that reversing a keyed item changes the coefficient in the expected direction. Confirm that the coefficient falls when the sample is restricted to a narrower range, since that is the signature of a population-dependent quantity behaving correctly. And before quoting any coefficient as evidence of quality, write the sentence saying what the scores are for; if that sentence cannot be written, the coefficient is not evidence of anything.
Application
Where the interpretation decides the answer
A reading test validated for identifying pupils who need support, now used to rank schools. The scores are unchanged and the validity evidence does not transfer, because the interpretation has changed. Evidence that individual scores identify pupils needing help says nothing about whether school means measure school quality, aggregation introduces intake composition, which the original argument never addressed.
A depression screen with
---
In the first three the coefficient is doing its job and is being read without the context that makes it interpretable. The number of items, the population. In the last two the coefficient is fine and the question being asked of it was never one it could answer.
---
Admissions and licensure testing. A cut score converts a continuous scale into a decision, so the standard error of measurement at that point is the operative quantity: candidates within about
Clinical screening. An instrument is validated for a purpose, case-finding in a stated population at a stated prevalence, and its performance does not transfer to a population with different base rates. Reliability is a precondition, since an unstable score cannot track anything, and the evidence that matters is how the scores behave against diagnosis in the population where the screen will be used.
Employee selection. Validity here is an argument connecting test content to job requirements, and it is contestable in court. What is defended is the interpretation, that these scores predict performance in this role, rather than the instrument, and a coefficient of internal consistency contributes almost nothing to that argument.
Survey scales in research. Published alphas are routinely quoted from the original development sample and applied to new populations. The coefficient should be recomputed on the data in hand, because it is a property of those scores; and where a scale has known facets, the subscale structure should be reported rather than a single total whose meaning depends on the facets being equivalent.
School accountability. Aggregating pupil scores to a school mean changes the interpretation entirely. Pupil-level validity evidence addresses individual attainment; school means additionally encode intake composition, so evidence for the first is not evidence for the second, however reliable the underlying scores.
---
The recurring decision. In each case the coefficient is available and answers a narrow question about consistency, while the decision rests on an interpretation that has to be argued and bounded. The discipline is to write down what the scores are taken to mean and for what decision, before quoting any number as evidence of quality.
Warning
Two errors about validity
And two errors about validity.
Treating an instrument as validated. Validity attaches to an interpretation of scores for a use. A reading test validated for identifying pupils needing support carries no evidence for ranking schools, because the interpretation changed while the scores did not. "A validated instrument" names no property unless the interpretation accompanies it.
Offering reliability as validity evidence. Reliability bounds validity: a score cannot correlate with a criterion more strongly than its own reliability permits. That makes it necessary and leaves it silent on whether the construct being measured consistently is the one the decision requires. A screen with