Subject
Psychometrics
Measuring things that cannot be observed directly, and saying what the resulting numbers support. An observed score is a true score plus error, reliability is the share of variance that is not error, and validity is not a further coefficient but an argument about a particular interpretation of scores for a particular use. The subject exists because those two questions are routinely collapsed into one, and a stable instrument can measure the wrong thing consistently.
Learning paths
Psychometrics
How a measurement of an unobservable construct is judged: what makes scores stable, and what would make an interpretation of them defensible.
What this subject develops
Say what a measurement supports before using it
Compute a reliability coefficient, report it on the score scale, and state the interpretation the scores are being put to and the evidence that would bear on it. The competency exists because the two questions asked of any measurement are routinely collapsed into one: whether the instrument gives a stable answer, and whether that answer means what it is taken to mean. A stable instrument can measure the wrong thing, and the coefficient most often quoted as proof of quality responds to test length as well as to item quality, so it can be raised without improving anything. Treating measurement as settled once a coefficient clears a threshold is the failure this competency is designed to prevent.
See detailed outcomes
Say what a measurement supports before using it
- Given a reliability coefficient and the score standard deviation, the learner can compute the standard error of measurement, state the band an individual score carries and the assumptions that band rests on, and attribute the coefficient to scores from a population rather than to the instrument.
- Given item and total-score variances, the learner can compute coefficient alpha, state the measurement model under which it estimates reliability and when the lower-bound reading is available, apply the Spearman-Brown relation to separate item quality from test length, and decline to read dimensionality off the coefficient.
- Given a proposed use of test scores, the learner can state the interpretation the use relies on, name the kinds of evidence that would bear on it, and explain why reliability bounds a validity coefficient without supplying validity and why evidence for one use does not transfer to another.