The Best Outcome Measure Is the One You Take Twice
A measure taken once is a description, not evidence. How to choose functional outcome measures by repeat cost, sensitivity to change, and payer recognition.
Outcome first
Functional goal builder
Activity
What will change?
Conditions
Where and with what support?
Measure
How will progress be visible?
Person + action + context + measure + time
At a glance
What you’ll leave with
- A baseline score with no second administration is a description, not evidence of change. The selection question is not "which measure is best" but "which measure will I actually repeat."
- Three criteria decide that: administration time (the cost you pay at every repeat), sensitivity to change (whether the score can move in your treatment window), and payer recognition (whether a reviewer can read the result as progress).
- Interpret change against the measure’s minimal detectable change (MDC) where one is published — a one-point shift inside the error band of the instrument is not progress a reviewer has to accept.
Most therapy charts contain a graveyard of outcome measures: a thorough evaluation with four standardized scores, none of which is ever administered again. The instrument was chosen for the evaluation — comprehensive, impressive, diagnostic — and it is exactly those qualities that guarantee nobody has forty minutes to repeat it on a Tuesday between sessions. A measure taken once is a description of a person on a day. It cannot show change, and showing change is the entire job. So the selection question for a working SLP, OT, or PT is not "which outcome measure is best for this condition?" It is "which measure am I certain I will take again in a few weeks, under the same conditions, in the time a real caseload allows?" This article gives you three criteria for answering that question, a comparison of the main classes of measures against them, and a cadence that makes the repeat administration automatic instead of aspirational.
The problem
A measure taken once proves nothing
Payers do not reimburse therapy because the patient is impaired; they reimburse it because skilled treatment is producing change the record can demonstrate. Medicare’s documentation manual makes the mechanism explicit: the chart needs objective measurement, and the progress report — due at least once every 10 treatment days under Part B — must present objective evidence of progress toward the goals. A baseline score with no follow-up administration cannot do that work. Neither can a second administration of a different instrument, because two different rulers cannot measure one distance. The evidence payers and referrers can actually use is the same measure, administered the same way, at two points in time.
This is also why the evaluation battery is usually the wrong place to look for your outcome measure. Norm-referenced diagnostic assessments answer a different question — how does this person compare to peers? — and answering it well makes them long, expensive to administer, and often explicitly not designed for frequent re-administration. The outcome measure is a different tool with a different job: small, repeatable, and pointed at the function you are treating. Choose it separately, and choose it at the evaluation, because the baseline you fail to take on day one cannot be reconstructed later.
The framework
Three criteria for choosing functional outcome measures
Score every candidate measure on three questions, in this order. First, administration time — not the time it takes once, but the cost you pay at every repeat. A measure you administer at baseline, at every progress report, and at discharge gets administered five or six times in a typical episode; a ten-minute difference per administration is an hour of clinical time over the episode, multiplied across the caseload. Second, sensitivity to change: can the score realistically move within your treatment window, for your population, in a way that exceeds the instrument’s own noise? A tool with a ceiling your patient will hit in three weeks, or one so coarse that a season of progress moves it one point, will produce a flat line that argues against your own treatment. Third, payer recognition: when the number appears in a progress report, can a reviewer who has never met the patient read it as functional progress? Measures anchored to observable daily function travel better in a chart than raw scores that need a manual to interpret.
Reading change honestly
Sensitivity to change: MDC is the number to know
Every instrument has measurement error, and researchers publish two thresholds that tell you how to read a score change against it. The minimal detectable change (MDC) is the smallest change larger than the instrument’s random error — movement below it may be noise, not progress. The minimal clinically important difference (MCID) is the smallest change patients experience as meaningful improvement. These are established per instrument and per population in the peer-reviewed literature, and they are the difference between "the score went up two points" and "the score improved beyond measurement error." Before adopting a measure, look up whether an MDC has been published for a population like yours; when you report change, report it against that threshold rather than celebrating movement the instrument itself cannot distinguish from noise. Where no threshold is published for your population, say what the change means in observed function instead of implying statistical certainty the literature does not offer.
The centerpiece
The five classes of measures, scored on all three criteria
Nearly every instrument you will consider falls into one of five classes, and the classes behave consistently on the three criteria even though individual instruments vary. Choose the class first, then the instrument inside it that has published measurement properties for your population.
Which class of measure survives repetition
Comparison| Class | Examples | Cost to repeat | Sensitivity to change | Payer readability |
|---|---|---|---|---|
| Performance-based tests | Timed, counted, or scored task performance — gait speed, timed transfers, structured task probes | Low — minutes, done inside the session with everyday equipment | Good for the specific function tested; watch for ceiling effects as the patient recovers | High — time and distance are self-explanatory to any reviewer |
| Patient- or caregiver-reported questionnaires | Region- or condition-specific scales the patient completes about daily function | Low for the clinician — the patient does the work in the waiting room or portal | Designed for it; published MDC values exist for many established scales | High when items describe daily activities; scores need the scale named and the direction explained |
| Individualized patient-specific scales | Instruments where the patient names their own activities and rates them, re-rated at intervals | Very low — a few ratings collected in conversation | High, because the items are the patient’s own limiting activities | Strong narrative fit, but reviewers may want a standardized companion measure alongside it |
| Clinician-rated functional scales | Ordinal ratings of functional level, such as ASHA’s seven-point Functional Communication Measures within NOMS | Very low — a judgment made from the session itself | Coarse steps: fine for episode-level change, weak for visit-to-visit movement | Good when the scale is nationally recognized; a home-made rating scale carries no weight |
| Norm-referenced diagnostic batteries | Full standardized assessments used at evaluation to establish eligibility and diagnosis | Very high — often an entire session, with strict administration rules | Poor for short intervals; many are not designed or normed for frequent re-administration | High for establishing baseline severity, poor for showing treatment-window progress |
The pattern in that table is the article’s argument in miniature: the classes that are cheap to repeat are the ones that show change, and the expensive one is the one designed for a different question. A defensible chart usually pairs two cheap classes — one standardized instrument for comparability, one individualized or performance-based measure for the functions this patient actually cares about — and leaves the diagnostic battery at the evaluation where it belongs.
By discipline
Where each discipline should start looking
Physical therapists have the most developed infrastructure for this decision. APTA maintains a searchable Tests and Measures library with summaries of individual instruments, and APTA Acute Care has published a clinical practice guideline recommending a core set of measures for assessing physical function in hospitalized adults — a useful model of the "small set, repeated often" philosophy even for outpatient settings. When a clinical practice guideline for your patient population names specific instruments, start there: guideline-named measures come with published measurement properties and instant recognizability.
Occupational therapists should anchor selection to the outcomes framework in AOTA’s Occupational Therapy Practice Framework: outcomes in OT are results in occupational performance and participation, not component skills alone. That argues for pairing one standardized instrument with an individualized, patient-specific measure in which the client names and rates the occupations that matter to them — the class of measure that is cheapest to repeat and speaks most directly to why the person came. A grip-strength number can support the story; the re-rated occupational performance score is the story.
Speech-language pathologists have the hardest version of the problem, because so much of the SLP toolbox is norm-referenced diagnostic assessment that resists frequent re-administration. Two workable paths: ASHA’s National Outcomes Measurement System (NOMS) rates function on seven-point Functional Communication Measures from evaluation to discharge, which gives episode-level change a nationally recognized vocabulary; and structured probe data — counted accuracy on defined targets under defined conditions — gives visit-level sensitivity that a seven-point scale cannot. The pairing covers both time horizons: probes move week to week, the FCM moves across the episode.
The system
A re-measurement cadence that runs itself
Selection fails in the calendar, not the catalog. The fix is to make re-measurement an appointment the schedule already contains rather than a task someone must remember. Under Medicare Part B the progress report is due at least every 10 treatment days and must contain objective evidence of progress — so the progress-report schedule is a re-measurement schedule the payer has already built for you. Other payers set their own intervals; the mechanic works the same way once you know the interval.
- 01
Take the baseline at the evaluation, and record the conditions
Administer the chosen measure on day one and note how: position, equipment, instructions, assistance level, time of day if it matters. A repeat under different conditions measures the conditions, not the patient.
- 02
Name the re-measurement trigger in the plan
Write the sentence at the evaluation: "Re-administer [measure] at each progress report and at discharge." Tie it to the reporting interval your payer already enforces so the deadline arrives with the paperwork that needs the number.
- 03
Repeat under the same conditions, inside a normal session
If the repeat cannot fit inside a treatment session, the measure failed criterion one — swap it for a shorter one now, at the first repeat, rather than quietly abandoning measurement for the rest of the episode.
- 04
Interpret against the published threshold
Compare the change to the instrument’s published MDC for a comparable population when one exists. Change beyond it is progress you can assert; change inside it is a reason to keep treating and re-measure at the next interval, stated honestly.
- 05
Put the comparison next to the goal, in one sentence
The progress report line a reviewer can use reads: baseline score, current score, threshold, and what the change means in daily function. Two administrations of one instrument, one sentence of interpretation.
- 06
At discharge, close the loop with the same instrument
The discharge administration turns the episode into a complete data story — baseline, trajectory, endpoint — that supports the discharge decision and gives the referrer a result they can read in ten seconds.
The criteria in use
Choosing and repeating a measure, start to finish
Worked example — fictional
An OT episode where the cheap measure carries the chart
A fictional case, built to show the selection logic rather than any real patient. An occupational therapist evaluates an adult recovering from a distal radius fracture whose stated problems are cooking dinner, typing a full workday, and carrying groceries.
The therapist considers a full standardized battery, a region-specific patient-reported questionnaire for the upper limb, and a patient-specific scale in which the client rates her own three named activities. The battery fails criterion one — it cannot be repeated inside a session. The questionnaire and the patient-specific ratings both pass all three criteria, so she keeps both: the questionnaire for standardized comparability, the self-rated activities because they are the reason the client came.
At the evaluation the client completes the questionnaire in the waiting room and rates her three activities in conversation. The therapist records the scores, the administration conditions, and one sentence in the plan: re-administer both at every progress report and at discharge. Total added evaluation time: about ten minutes, most of it the client’s.
At the first progress report both measures are repeated the same way. The questionnaire score has improved by more than the change threshold the therapist found in the published literature for this instrument; typing has moved from "cannot sustain ten minutes" toward a rating consistent with a half day. The progress-report line writes itself: instrument, baseline, current score, threshold, and what changed at the desk and in the kitchen.
A reviewer who has never met the client sees the same instrument at two time points, movement beyond the instrument’s error band, and function stated in everyday terms. Nothing in the chart depends on the reviewer trusting the therapist’s adjectives — the argument is carried by two administrations of the same ruler.
Criterion three, expanded
What payer recognition actually means
Payer recognition does not mean there is a national list of approved instruments — for most payers there is not, and Medicare’s manual asks for objective measurement without prescribing a catalog. Some Medicare contractors and commercial plans do name preferred instruments in their local coverage policies and medical-necessity guidelines, and those documents, not habit, are the authority for a specific payer. Recognition in practice is humbler: the measure is named, the score has a direction a reader can follow, the same measure appears at least twice, and the change is translated into function. A chart can fail all four with a famous instrument and pass all four with structured probe data, so spend your care on the repetition and the translation rather than on hunting for a magic instrument.
Quick answers
Functional outcome measures: the questions that keep coming up
What are functional outcome measures in therapy?
Standardized or structured tools that quantify how a person performs meaningful activities — walking, dressing, communicating, swallowing, working — so that change over an episode of care can be demonstrated rather than described. They differ from diagnostic assessments, which compare a person to peers to establish the presence and severity of a condition. The same patient usually needs both, but only the outcome measure has to be repeated.
How often should outcome measures be repeated?
At minimum: baseline at the evaluation, again at each required progress reporting interval, and at discharge. Under Medicare Part B the progress report is due at least once every 10 treatment days and must include objective evidence of progress, which makes that interval a natural re-measurement schedule. Commercial payers, Medicaid programs, and school systems set their own intervals — check the contract or coverage policy rather than assuming Medicare’s.
What is the difference between MDC and MCID?
The minimal detectable change (MDC) is the smallest score change larger than the instrument’s measurement error — below it, the change may be noise. The minimal clinically important difference (MCID) is the smallest change patients perceive as meaningful benefit. Both are estimated per instrument and per population in peer-reviewed studies. For defending progress in documentation, MDC is the workhorse: change beyond it is change the instrument can actually vouch for.
Do I still have to report G-codes to Medicare?
No. CMS discontinued functional limitation reporting — the nonpayable G-codes and severity modifiers — for outpatient therapy claims with dates of service on and after January 1, 2019. The underlying documentation expectation survives: evaluations need objective measurement and progress reports need objective evidence of progress toward the goals.
Can I use a norm-referenced assessment as my outcome measure?
Usually not well. Norm-referenced batteries are built to compare a person to peers at a point in time; many are long, and some are not designed or normed for frequent re-administration. They earn their place at the evaluation and at major decision points. For showing change inside a treatment episode, pair the battery with something cheap to repeat — a performance-based test, a patient-reported scale, or structured probe data.
What if my patient improved but the score change is smaller than the MDC?
Say exactly that, honestly: the observed change is within the instrument’s error band, here is what changed in observed function, and here is the plan — keep treating and re-measure at the next interval, add a more sensitive measure, or reconsider the approach. What undermines a chart is claiming a one-point move as proof; what supports it is a clinician who visibly knows what the instrument can and cannot show.
Primary sources
Bibliography / 9- 01Medicare Benefit Policy Manual, Pub. 100-02, Chapter 15, Section 220.3 (documentation requirements for therapy services)Centers for Medicare & Medicaid Services
- 02Functional Reporting (discontinuation of therapy G-codes effective January 1, 2019)Centers for Medicare & Medicaid Services
- 03Medicare Part B Documentation RequirementsAmerican Physical Therapy Association
- 04Tests and Measures (evidence-based instrument summaries)American Physical Therapy Association
- 05A Core Set of Outcome Measures to Assess Physical Function for Adults (clinical practice guideline)American Physical Therapy Association, Academy of Acute Care Physical Therapy
- 06National Outcomes Measurement System (NOMS) and Functional Communication MeasuresAmerican Speech-Language-Hearing Association
- 07Occupational Therapy Practice Framework: Domain and Process — Intervention OutcomesAmerican Occupational Therapy Association
- 08Determining the Minimal Clinically Important Difference for Rehabilitation Measures (research review)Model Systems Knowledge Translation Center (NIDILRR, Administration for Community Living)
- 09Assessing the Stroke-Specific Quality of Life for Outcome Measurement in Stroke Rehabilitation: Minimal Detectable Change and Clinically Important Difference (peer-reviewed example of MDC/MCID estimation)Health and Quality of Life Outcomes, via PubMed Central
Written by Callie Editorial
Published September 2, 2026
Educational content, not legal, billing, or patient-specific clinical advice.
Talk to our team