Validity as an Argument for an Interpretation
Validity as the degree to which evidence supports a specific interpretation of scores for a specific use, the kinds of evidence that bear on such a claim, why a consistent score can be consistently measuring the wrong thing, and why evidence gathered for one use does not transfer to another.
Definition
Validity is not a further coefficient. It is the degree to which evidence supports a specific interpretation of scores for a specific use, so the same instrument can be validly interpreted one way and not another, and a validity claim names the interpretation before the evidence.
Assumptions and scope
Validity attaches to an interpretation for a use. A test validated for one purpose carries no validity evidence for another, and the phrase "a validated instrument" names no property unless the interpretation is stated.
Reliability limits validity coefficients without supplying validity. Under the classical model with uncorrelated errors, an observed correlation is attenuated from the true-score correlation by
, so unreliability in either measure caps the observed association. A perfectly reliable score may still correlate with nothing of interest.
Worked material
Example
One instrument, four proposed uses
A 25-item numeracy assessment,
---
Use 1: telling a teacher which pupils need extra practice. The interpretation is that a low score indicates a pupil who would benefit from additional instruction now. Evidence bearing on it: whether low scorers do improve more than high scorers when given the instruction, and whether the items cover what the instruction addresses. The
Use 2: certifying that a pupil has met a standard. The interpretation is that a score above a cut represents mastery of a defined content domain. Now content evidence becomes central: does the item set sample the domain the standard names, in the proportions the standard implies? A correlation with later attainment is close to irrelevant here, and the operative statistic is classification consistency near the cut, computed from the standard error of measurement rather than from the coefficient.
Use 3: comparing this year's cohort against last year's. The interpretation is that a difference in means reflects a difference in numeracy rather than in the measurement. This requires evidence the two administrations were equivalent: the same items, or equated forms, and no change in who sits the assessment. An instrument perfectly valid for Use 1 supports this only if that equivalence holds, and nothing in the reliability coefficient speaks to it.
Use 4: evaluating teachers by their class's mean score. The interpretation is that class mean reflects teaching quality. This requires everything above plus an argument that intake composition has been accounted for, since a class mean encodes who is in the class as much as what happened in it. Consequential evidence also becomes central: once scores determine evaluations, teaching narrows toward the tested items, which changes what the score measures.
---
| Use | Evidence that matters most | What |
|---|---|---|
| Identify pupils to support | response to instruction | necessary precondition only |
| Certify a standard | content coverage, classification consistency | necessary precondition only |
| Compare cohorts | administration equivalence | necessary precondition only |
| Evaluate teachers | intake adjustment, consequences | necessary precondition only |
The column on the right is the point. A single reliability coefficient contributes the same thing to all four, and it is never the thing in dispute. What differs down the left column is what would have to be shown, and each row is a separate argument that has to be made rather than inherited from the row above.
And the direction of failure. Uses 1 and 2 are the ones the instrument was built for. Use 4 is the furthest from it, and it is the one most likely to be adopted without a new argument, because the scores are already there and already look authoritative.
Common errors
Common misconception
That validity is a property an instrument has, so that a test can be described as validated once and used thereafter for whatever purpose. Validity is the degree to which evidence supports a particular interpretation of scores for a particular use, which means the same instrument can be validly interpreted one way and invalidly another. A reading test validated for identifying pupils needing support carries no evidence for ranking schools, because the interpretation has changed even though the scores have not. The phrase "a validated instrument" names no property unless the interpretation and the use accompany it, and reliability does not supply the missing evidence: a score can be highly stable and mean nothing relevant to the decision it is informing.
Related units
Requires
Learn this topic
Used in
Sources
- Construct Validity in Psychological Tests (1955)
- Standards for Educational and Psychological Testing (2014)