reliability, validity, responsiveness, and minimal clinically important difference
/ MCID /
A measurement tool is only worth using if you can trust it, and trusting it means asking several pointed questions. Does it give the same answer when nothing has really changed? Does it actually measure what it claims to? Can it detect real improvement when it happens? And if the score does change, how big a change is big enough to matter to the person? These four questions are the psychometric properties of an outcome measure — the quality checks that separate a meaningful number from a misleading one.
Reliability asks whether the tool is consistent — whether two examiners (inter-rater) or the same examiner twice (test-retest) get close to the same result when the person has not changed; a tool that gives wildly different scores by chance is like a bathroom scale that reads differently each time you step on. Validity asks whether the tool measures what it intends — a balance test that mostly reflects leg strength is not really measuring balance. Responsiveness (also called sensitivity to change) asks whether the tool can pick up genuine improvement or decline rather than sitting flat while the person changes. And the minimal clinically important difference, or MCID, is the smallest change in score that a person would actually notice and value — a one-point change might be statistically real yet too small for anyone to feel.
These properties matter because rehabilitation runs on numbers, and a number with poor properties can do harm: a tool with weak reliability makes you chase noise, one with weak validity measures the wrong thing confidently, and a change smaller than the MCID can be celebrated as success when the person feels nothing different. Knowing them lets a clinician choose the right tool for the question, set honest goals, and read a research study critically. The key humility is that these properties are not fixed labels but depend on the setting and the population — a measure that is reliable and responsive in stroke may behave quite differently in another condition.
A study reports that a new therapy improved a balance score by 2 points, which is statistically significant. But the test's minimal clinically important difference is about 5 points, so the change, though real, is too small for patients to feel — a result that looks like success on paper but means little in life.
Statistically significant and clinically meaningful are not the same thing.
These properties are not permanent labels: a measure that is reliable, valid, and responsive in one condition or setting can behave differently in another, so good practice is to ask whether a tool was validated in people like the one in front of you.