Validity as an Argument for an Interpretation

Validity as the degree to which evidence supports a specific interpretation of scores for a specific use, the kinds of evidence that bear on such a claim, why a consistent score can be consistently measuring the wrong thing, and why evidence gathered for one use does not transfer to another.

Definition

Validity is not a further coefficient. It is the degree to which evidence supports a specific interpretation of scores for a specific use, so the same instrument can be validly interpreted one way and not another, and a validity claim names the interpretation before the evidence.

Assumptions and scope

  • Validity attaches to an interpretation for a use. A test validated for one purpose carries no validity evidence for another, and the phrase "a validated instrument" names no property unless the interpretation is stated.

  • Reliability limits validity coefficients without supplying validity. Under the classical model with uncorrelated errors, an observed correlation is attenuated from the true-score correlation by ρ X Y = ρ T X T Y ρ X X ′ ρ Y Y ′ , so unreliability in either measure caps the observed association. A perfectly reliable score may still correlate with nothing of interest.

Worked material

Example

One instrument, four proposed uses

A 25-item numeracy assessment, α = 0.89 , correlating 0.44 with end-of-year mathematics attainment. The scores never change below. The interpretation does, and with it what must be argued.

---

Use 1: telling a teacher which pupils need extra practice. The interpretation is that a low score indicates a pupil who would benefit from additional instruction now. Evidence bearing on it: whether low scorers do improve more than high scorers when given the instruction, and whether the items cover what the instruction addresses. The 0.44 is weak support, since it concerns attainment rather than response to teaching.

Use 2: certifying that a pupil has met a standard. The interpretation is that a score above a cut represents mastery of a defined content domain. Now content evidence becomes central: does the item set sample the domain the standard names, in the proportions the standard implies? A correlation with later attainment is close to irrelevant here, and the operative statistic is classification consistency near the cut, computed from the standard error of measurement rather than from the coefficient.

Use 3: comparing this year's cohort against last year's. The interpretation is that a difference in means reflects a difference in numeracy rather than in the measurement. This requires evidence the two administrations were equivalent: the same items, or equated forms, and no change in who sits the assessment. An instrument perfectly valid for Use 1 supports this only if that equivalence holds, and nothing in the reliability coefficient speaks to it.

Use 4: evaluating teachers by their class's mean score. The interpretation is that class mean reflects teaching quality. This requires everything above plus an argument that intake composition has been accounted for, since a class mean encodes who is in the class as much as what happened in it. Consequential evidence also becomes central: once scores determine evaluations, teaching narrows toward the tested items, which changes what the score measures.

---

UseEvidence that matters mostWhat α = 0.89 contributes
Identify pupils to supportresponse to instructionnecessary precondition only
Certify a standardcontent coverage, classification consistencynecessary precondition only
Compare cohortsadministration equivalencenecessary precondition only
Evaluate teachersintake adjustment, consequencesnecessary precondition only

The column on the right is the point. A single reliability coefficient contributes the same thing to all four, and it is never the thing in dispute. What differs down the left column is what would have to be shown, and each row is a separate argument that has to be made rather than inherited from the row above.

And the direction of failure. Uses 1 and 2 are the ones the instrument was built for. Use 4 is the furthest from it, and it is the one most likely to be adopted without a new argument, because the scores are already there and already look authoritative.

Common errors

Common misconception

That validity is a property an instrument has, so that a test can be described as validated once and used thereafter for whatever purpose. Validity is the degree to which evidence supports a particular interpretation of scores for a particular use, which means the same instrument can be validly interpreted one way and invalidly another. A reading test validated for identifying pupils needing support carries no evidence for ranking schools, because the interpretation has changed even though the scores have not. The phrase "a validated instrument" names no property unless the interpretation and the use accompany it, and reliability does not supply the missing evidence: a score can be highly stable and mean nothing relevant to the decision it is informing.

Related units

Requires

Learn this topic

Used in

Sources

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.