Validity as an Argument for an Interpretation

What you will be able to do

Given a proposed use of test scores, the learner can state the interpretation the use relies on, name the kinds of evidence that would bear on it, and explain why reliability bounds a validity coefficient without supplying validity and why evidence for one use does not transfer to another.

Orientation

A claim about an interpretation, not about an instrument

Reliability has a number attached. Validity does not, and that difference is structural rather than a gap in the literature.

Validity is the degree to which evidence supports a specific interpretation of scores for a specific use. So the same instrument can be validly interpreted one way and not another, and the phrase "a validated instrument" names no property until the interpretation is stated.

Two consequences follow, and both are routinely violated in practice.

Evidence does not transfer across uses. A reading test validated for identifying pupils who need support carries no validity evidence for ranking schools. The scores are unchanged; the interpretation changed, and the original argument never addressed the new one.

Reliability is necessary and nowhere near sufficient. A score cannot correlate with a criterion more strongly than its own reliability permits, so consistency bounds validity. It supplies none of it: a depression screen with α = 0.94 correlating 0.12 with clinical diagnosis satisfies the bound and fails the purpose completely.

This unit covers what a validity argument consists of, what kinds of evidence bear on it, and why the argument has to name the decision before it cites a number.

Definition

What a validity claim is defined over

Validity is defined over an interpretation and a use, not over the instrument. The standard formulation asks how far evidence supports a specified interpretation of scores for a specified purpose. Two consequences follow. A test carries no validity in the abstract, so "a validated instrument" names nothing until the interpretation is stated. And evidence gathered for one use does not transfer to another, however similar the scores look.

Reliability bounds validity without supplying it. A score cannot correlate with an external criterion more strongly than its own reliability permits, which makes reliability necessary. It is not sufficient: a perfectly reliable score may be reliably measuring something irrelevant to the decision at hand.

Intuition

What settles a validity claim

Validity is a different kind of claim. Nothing above bears on whether the scores mean what they are being used to mean. That question is settled by stating the interpretation, identifying what would have to hold for it to be sound, and gathering evidence on each, including evidence that could count against it.

Example

One instrument, four proposed uses

A 25-item numeracy assessment, α = 0.89 , correlating 0.44 with end-of-year mathematics attainment. The scores never change below. The interpretation does, and with it what must be argued.

---

Use 1: telling a teacher which pupils need extra practice. The interpretation is that a low score indicates a pupil who would benefit from additional instruction now. Evidence bearing on it: whether low scorers do improve more than high scorers when given the instruction, and whether the items cover what the instruction addresses. The 0.44 is weak support, since it concerns attainment rather than response to teaching.

Use 2: certifying that a pupil has met a standard. The interpretation is that a score above a cut represents mastery of a defined content domain. Now content evidence becomes central: does the item set sample the domain the standard names, in the proportions the standard implies? A correlation with later attainment is close to irrelevant here, and the operative statistic is classification consistency near the cut, computed from the standard error of measurement rather than from the coefficient.

Use 3: comparing this year's cohort against last year's. The interpretation is that a difference in means reflects a difference in numeracy rather than in the measurement. This requires evidence the two administrations were equivalent: the same items, or equated forms, and no change in who sits the assessment. An instrument perfectly valid for Use 1 supports this only if that equivalence holds, and nothing in the reliability coefficient speaks to it.

Use 4: evaluating teachers by their class's mean score. The interpretation is that class mean reflects teaching quality. This requires everything above plus an argument that intake composition has been accounted for, since a class mean encodes who is in the class as much as what happened in it. Consequential evidence also becomes central: once scores determine evaluations, teaching narrows toward the tested items, which changes what the score measures.

---

UseEvidence that matters mostWhat α = 0.89 contributes
Identify pupils to supportresponse to instructionnecessary precondition only
Certify a standardcontent coverage, classification consistencynecessary precondition only
Compare cohortsadministration equivalencenecessary precondition only
Evaluate teachersintake adjustment, consequencesnecessary precondition only

The column on the right is the point. A single reliability coefficient contributes the same thing to all four, and it is never the thing in dispute. What differs down the left column is what would have to be shown, and each row is a separate argument that has to be made rather than inherited from the row above.

And the direction of failure. Uses 1 and 2 are the ones the instrument was built for. Use 4 is the furthest from it, and it is the one most likely to be adopted without a new argument, because the scores are already there and already look authoritative.

Procedure

Assembling a validity argument

To assemble a validity argument.

  1. State the interpretation and the use in one sentence, before any evidence: what the scores are taken to mean, and what decision they will inform.
  2. List what must hold for that interpretation to be sound, relations with other measures, the content the items cover, the response processes respondents actually use, the consequences of the decision.
  3. Gather evidence on each, including evidence that could count against the interpretation.
  4. State the boundary. Evidence supports this interpretation for this use; a different use needs its own argument.

Checks. Confirm that reversing a keyed item changes the coefficient in the expected direction. Confirm that the coefficient falls when the sample is restricted to a narrower range, since that is the signature of a population-dependent quantity behaving correctly. And before quoting any coefficient as evidence of quality, write the sentence saying what the scores are for; if that sentence cannot be written, the coefficient is not evidence of anything.

Application

Where the interpretation decides the answer

A reading test validated for identifying pupils who need support, now used to rank schools. The scores are unchanged and the validity evidence does not transfer, because the interpretation has changed. Evidence that individual scores identify pupils needing help says nothing about whether school means measure school quality, aggregation introduces intake composition, which the original argument never addressed.

A depression screen with α = 0.94 and a correlation of 0.12 with clinical diagnosis. Highly consistent and nearly useless for the purpose named. This is the clearest form of the point: reliability bounds validity, so a score cannot correlate with a criterion more strongly than its own reliability allows, and a reliable score may still be reliably measuring something else.

---

In the first three the coefficient is doing its job and is being read without the context that makes it interpretable. The number of items, the population. In the last two the coefficient is fine and the question being asked of it was never one it could answer.

---

Admissions and licensure testing. A cut score converts a continuous scale into a decision, so the standard error of measurement at that point is the operative quantity: candidates within about 2 SEM of the boundary are not separated by the instrument. Testing bodies publish the SEM and the classification consistency near the cut for this reason, and the overall coefficient is the least informative number in the report.

Clinical screening. An instrument is validated for a purpose, case-finding in a stated population at a stated prevalence, and its performance does not transfer to a population with different base rates. Reliability is a precondition, since an unstable score cannot track anything, and the evidence that matters is how the scores behave against diagnosis in the population where the screen will be used.

Employee selection. Validity here is an argument connecting test content to job requirements, and it is contestable in court. What is defended is the interpretation, that these scores predict performance in this role, rather than the instrument, and a coefficient of internal consistency contributes almost nothing to that argument.

Survey scales in research. Published alphas are routinely quoted from the original development sample and applied to new populations. The coefficient should be recomputed on the data in hand, because it is a property of those scores; and where a scale has known facets, the subscale structure should be reported rather than a single total whose meaning depends on the facets being equivalent.

School accountability. Aggregating pupil scores to a school mean changes the interpretation entirely. Pupil-level validity evidence addresses individual attainment; school means additionally encode intake composition, so evidence for the first is not evidence for the second, however reliable the underlying scores.

---

The recurring decision. In each case the coefficient is available and answers a narrow question about consistency, while the decision rests on an interpretation that has to be argued and bounded. The discipline is to write down what the scores are taken to mean and for what decision, before quoting any number as evidence of quality.

Warning

Two errors about validity

And two errors about validity.

Treating an instrument as validated. Validity attaches to an interpretation of scores for a use. A reading test validated for identifying pupils needing support carries no evidence for ranking schools, because the interpretation changed while the scores did not. "A validated instrument" names no property unless the interpretation accompanies it.

Offering reliability as validity evidence. Reliability bounds validity: a score cannot correlate with a criterion more strongly than its own reliability permits. That makes it necessary and leaves it silent on whether the construct being measured consistently is the one the decision requires. A screen with α = 0.94 correlating 0.12 with diagnosis satisfies the bound and fails the purpose.

Next step

Practice Validity as an Argument for an Interpretation

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.