Practice: Validity as an Argument for an Interpretation

Error diagnosis · Classification

A vendor writes: "Our 30-item screening instrument has α = 0.94 and has been validated. It is therefore suitable for deciding which employees receive development funding."

Which objection is the sharpest?

2 hints available, least help first.

Hint 1: Retrieval cue

What does a validity claim have to name before any evidence is cited?

Hint 2: Concept cue

Ask what question the coefficient answers, and whether it is the question in dispute.

Transfer · Evaluation

A benchmark of multiple-choice questions was built to compare language models on factual recall, and its scores proved highly stable across repeated runs and paraphrased prompts. A team now cites benchmark rank as evidence that a model is suitable for clinical decision support. Which objection is the sharpest?

Evaluation · Explanation

A university reports that its admissions test has α = 0.91 on last year's applicants and correlates 0.28 with first-year grade average. It proposes using the score to rank applicants for a scholarship.

(a) State the interpretation the proposal relies on, in one sentence, naming the decision it informs.

(b) Name three kinds of evidence that would bear on that interpretation, including one that could count against it.

(c) The committee argues the α = 0.91 shows the test is "a valid measure". Say precisely what that coefficient does and does not establish here, and what the 0.28 contributes.

Write your answer, then compare it with the worked solution.

3 hints available, least help first.

Hint 1: Retrieval cue

Write the interpretation as a sentence naming both what the score means and what decision it informs.

Hint 2: Concept cue

Ask what each reported number is computed from, and over which population.

Hint 3: Strategy cue

For (b), one kind of evidence should be capable of undermining the proposal, not only supporting it.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) The interpretation. The proposal relies on: this score measures academic preparation well enough that a higher score identifies an applicant more likely to benefit from, and succeed with, the scholarship. That is a claim about individual-level prediction, for a selection decision with a fixed number of awards.

Naming the decision matters, because the same test might be defensible for advising applicants about preparation and indefensible for allocating money between them.

(b) Evidence that would bear on it. Three kinds, the third of which could count against:

  • Criterion evidence at the individual level within the relevant range. The 0.28 is computed across all applicants; the decision concerns the narrow band competing for the scholarship, where range restriction will attenuate the relation further. The relevant question is how scores relate to outcomes among applicants near the cut.
  • Response-process evidence. Whether applicants answer from the preparation the test intends to measure, or from coaching, familiarity with the format, or test-taking speed. Evidence here would come from think-aloud studies or from comparing coached and uncoached groups.
  • Consequential evidence, which could count against the interpretation. Once a scholarship depends on the score, preparation behaviour changes, and any group difference in access to coaching becomes a difference in awards. If the score's relation to outcomes differs across groups, or if the ranking systematically disadvantages a group without corresponding differences in outcomes, that is evidence against the proposed use even if the correlation holds overall.

(c) What each number establishes. α = 0.91 says the items covary substantially in last year's applicant pool, under a measurement model with errors uncorrelated across items. It establishes that the score is stable enough to be worth interpreting, and nothing at all about what is being measured: a highly consistent test of coaching familiarity would report the same figure. "A valid measure" is not a property the coefficient can confer.

The 0.28 is the only figure here bearing on the interpretation, and it is weak: it accounts for under 8% of variance in first-year grades, it is computed on the full range rather than near the decision boundary, and it is an association rather than evidence that the scholarship will produce better outcomes for higher scorers. Note also that reliability bounds it. With ρ X X ′ = 0.91 and a perfectly reliable criterion the observed correlation could reach about 0.95 , so unreliability is not what limits this one. The construct relation is simply weak, and the proposal needs evidence it does not yet have.

A complete answer does each of these:

  • frames validity as interpretation

Construction · Direct application · Explanation

A hospital trust has a 20-item wellbeing questionnaire, α = 0.92 , developed and validated to help clinicians decide which outpatients to offer a follow-up appointment. It now proposes using ward-level mean scores to rank wards for a quality programme.

(a) State, in one sentence each, the interpretation the original use relies on and the interpretation the new use relies on.

(b) Say why the original evidence does not transfer, naming what specifically changed.

(c) Name three kinds of evidence bearing on the new interpretation, including one that could count against it.

(d) The trust argues α = 0.92 shows the instrument is 'validated'. Say what that coefficient establishes here and what it cannot.

Write your answer, then compare it with the worked solution.

3 hints available, least help first.

Hint 1: Retrieval cue

Write each interpretation as a sentence naming both the meaning and the decision.

Hint 2: Concept cue

In (b), ask what unit each interpretation is about.

Hint 3: Strategy cue

For (c), one kind of evidence should be able to undermine the proposal.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) The two interpretations.

Original: this patient's score indicates their current wellbeing well enough that a low score identifies someone who would benefit from a follow-up appointment.

Proposed: the mean score of a ward's patients indicates the quality of care that ward provides, well enough to rank wards against one another.

Writing both as sentences is what makes the difference visible. They share an instrument and share nothing else.

(b) Why the evidence does not transfer. Three things changed at once.

The unit changed, from a patient to a ward. A ward mean is not a larger version of a patient score; it aggregates over patients who were not sampled to represent the ward.

The construct changed, from an individual's wellbeing to an organisation's quality of care. A ward mean encodes case mix as much as care: a ward admitting sicker patients scores lower with no difference in what it does.

The decision changed, from offering one person an appointment to ranking units with consequences attached. Validity attaches to an interpretation for a use, so a change in any of the three requires its own argument. Nothing about the scores changed, which is exactly why the instrument feels transferable and is not.

(c) Evidence that would bear on it.

  • Case-mix adjustment. Whether ward differences survive adjustment for admission severity, age and diagnosis. Without it the ranking may be measuring intake.
  • Stability at the ward level. Whether a ward's mean is stable across months and across patient samples. Individual-level reliability says nothing about this: a highly reliable individual score can still produce an unstable mean from few respondents.
  • Consequential evidence, which could count against it. What happens once rankings carry consequences: whether response rates fall on low-ranked wards, whether favourable respondents are encouraged, and whether measured improvement is matched by an independent indicator. Improvement on the ranking alone is evidence against the interpretation.

(d) What the coefficient establishes. α = 0.92 says the 20 items covary substantially in the development sample, under the classical model with errors uncorrelated across items. It establishes that individual scores are consistent enough to be worth interpreting.

It cannot establish what is measured — a consistent questionnaire about catering would report a similar figure — and says nothing about the ward level, having been computed on individual responses. 'Validated' is not a property a coefficient confers, and names nothing until the interpretation and the use accompany it.

A complete answer does each of these:

  • frames validity as interpretation
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.