Skip to main content
The Practice
Clinical operationsAugust 3, 2026

The Review Pass That Makes an AI-Drafted Note Yours

AI-drafted therapy notes fail in ways tired human drafts do not: invented objective data, softened skilled rationale, and copied-forward assessments — all of it fluent and confident. This article gives SLPs, OTs, and PTs the specific review pass to run before signing, a worked example of the catches, and the checklist that keeps the pass fast.

Callie Editorial 15 min read
The sign-off issue
Your name
Session note
S

Relevant report

Caregiver reports carryover at home

O

Observable change

78% accuracy · minimal verbal cue

A

Clinical meaning

Self-monitoring is emerging

P

Next decision

Progress to conversational retell

At a glance

What you’ll leave with

  • Medicare requires the author of a service to authenticate it, and CMS told reviewers years ago — for human scribes — that the treating clinician’s signature is the one that matters. Signing an AI draft makes you its author, with no asterisk.
  • AI drafts fail differently from tired human drafts: instead of gaps, they produce fluent, confident errors. Three patterns account for most of them — invented objective data, skilled rationale softened into bystander language, and assessments copied forward from the last note.
  • A fixed sign-off checklist, run in the same order every time, turns review from passive re-reading into targeted hunting — and every recurring catch is a configuration fix that belongs in the tool’s template, not just in today’s note.

The pitch delivered: the session ended, and a draft note was waiting. It reads clean. The objective section has numbers, the assessment sounds clinical, the plan is plausible. That polish is exactly the problem. When a tired clinician writes a weak note, the weakness is visible — missing data, a thin assessment, a sentence that trails off. When a language model writes a weak note, the weakness is fluent: a percentage that was never tallied, skilled reasoning flattened into activity description, yesterday’s assessment wearing today’s date. The moment you sign, none of it is the software’s draft anymore. It is your clinical record, your billing support, and your testimony if anyone ever asks. This article is about the review pass that happens between the draft and the signature — what specifically to hunt for, in what order, and how to keep the pass fast enough that you actually do it every time.

The stakes

Your signature makes you the author

Medicare’s rule is old, short, and indifferent to how the note was produced: services must be authenticated by the author, by handwritten or electronic signature. And CMS has already answered the question of what happens when someone else drafts the note. In its 2017 Program Integrity Manual update on scribe services, CMS instructed reviewers to look only for the treating clinician’s signature — the scribe’s does not matter — because the clinician’s signature affirms that the note adequately documents the care provided. Swap the human scribe for a model and the logic does not move an inch. There is no lower standard for AI-assisted notes, no shared-blame arrangement with the vendor, and no reviewer anywhere who will weigh “the software wrote that part” as a defense. The draft belongs to the tool; the signed note belongs to you.

The professional associations have converged on the same position from the clinical side. APTA’s practice advisory on AI-enabled ambient scribe technology is direct: AI-generated clinical documentation may contain errors of omission or addition, or nuanced inaccuracies, and must always be reviewed. ASHA’s guidance on generative AI for clinicians frames these tools as support for clinical work, not a substitute for clinical reasoning, and points to the Code of Ethics obligation to evaluate any technology used in practice. AOTA’s journal has made the same call for critical evaluation of AI tools in occupational therapy practice. None of this is anti-AI — all three associations acknowledge the documentation burden these tools attack. It is a statement about where the review responsibility lands, and it lands on the signer.

Know your enemy

The three failure modes worth hunting

A review pass that just re-reads the note will miss what matters, because AI drafts do not fail the way human drafts fail. A rushed human note has visible gaps. An AI draft has invisible substitutions — content that is grammatical, plausible, and wrong in ways that only the person who was in the room can detect. Three patterns account for most of them, and each one concentrates in a different section of the note.

1. Invented objective data

Language models complete patterns, and therapy notes are full of patterns that look like “producing target with 80% accuracy across 20 trials.” Run a session where you cued heavily but never tallied, and the draft may hand you a percentage anyway — not because the tool heard one, but because notes like yours usually contain one. The same mechanism invents trial counts, standardized score fragments, distances, and minutes. Transcription adds its own class of errors: left becomes right, the target sound migrates, a parent’s report in the waiting room becomes something you observed. The objective section deserves the most suspicion precisely because it is the section reviewers trust most, and because a fabricated number is the one error that reads as completely normal on a skim. The fix is a hard rule, not vigilance in general: every number in the draft is either data you collected today or it comes out. Deleting is not always enough — if no tally was taken, say what you observed and that formal data was not collected, which is an honest note rather than a hole.

2. Skilled rationale softened into bystander language

Medicare’s documentation standard for therapy turns on whether the record shows that the service required the skills of a therapist — that is the load-bearing wall in the Benefit Policy Manual’s therapy documentation requirements, and most payer review follows the same logic. A general-purpose summarizer does not know that wall exists. It describes what happened in the room the way a bystander would: “patient practiced word retrieval with support,” “completed transfer training with assistance.” Nothing in those sentences is false. They simply describe activity anyone could have supervised, which is exactly the reading a reviewer is looking for permission to make. Your session contained decisions — the cue level you chose and faded, the task you graded down and why, the compensation you disallowed because it would block the goal. The draft will reliably sand those decisions into “support” and “assistance.” The review question for the assessment and intervention lines is not “is this accurate?” but “could an untrained person have done what this sentence describes?” Wherever the answer is yes and the truth is no, rewrite the line to name the decision.

3. Copied-forward assessments

Tools that read your previous note for context inherit its gravity: yesterday’s assessment language drifts into today’s draft, sometimes verbatim, and a caseload documented by the same tool converges on the same sentences across different patients. This is the oldest documentation risk in electronic records wearing new machinery. Medicare’s review contractors have warned for years that entries worded exactly like or similar to previous entries are treated as cloned documentation that fails to support medical necessity, and HHS’s Office of Inspector General flagged copy-paste misuse as a fraud vulnerability back in 2013 — when it found that only about a quarter of hospitals even had a policy governing the feature. An AI drafting layer can reproduce that pattern at a scale no copy-paste habit ever reached. The tell is an assessment that would have been equally true last week. The defense is to ask, before signing: what does this assessment say that the previous note could not have said? If the answer is nothing, today’s skilled observation is missing — and a reviewer reading the chart in sequence will notice, because reading in sequence is what reviewers do.

Worked example

One draft, three catches

A fictional pediatric speech session, invented for illustration. The clinician ran /r/ articulation drills and a story-retell task. Mid-session, she dropped the planned minimal-pairs activity because the child was fatiguing, and she faded cues from tactile placement to verbal-only. No formal tally was collected.

What the draft said

“Client produced /r/ in the initial position with 85% accuracy across 20 trials. Client remains motivated and engaged in therapy activities. Continue current treatment plan.”

What the review caught

All three failure modes in three sentences. The 85% figure was invented — no tally was taken. “Remains motivated and engaged” had appeared, word for word, in the last four notes. And nothing recorded the mid-session clinical decision: the dropped activity, the reason, or the cue fading that was the actual skilled work of the session.

What got signed

“Clinician provided tactile placement cues for initial /r/, faded to verbal-only cues by the third drill set; formal accuracy data not collected this session — clinician judgment: productions more consistent at single-word level than in the prior session. Discontinued planned minimal-pairs task due to fatigue; substituted story-retell to maintain production practice in connected speech. Next session: begin at verbal-only cue level and collect baseline tally at word level.”

The centerpiece

The sign-off checklist

Run the items in order, every time — the order goes where the risk is, numbers first. A fixed sequence is what makes the pass fast: you stop re-reading the note as prose and start hunting for specific defects, the way you scan an EOB or a schedule. On a routine session note, most items take seconds. The pass earns its keep on the items that occasionally take longer, because those are the catches that matter.

Field checklist

12 items

The AI-drafted note sign-off checklist

  • Every number is mine: each percentage, trial count, score, distance, and minute figure in the draft is data I actually collected today — anything else is deleted or replaced with what I observed, stated as observation.
  • Time reconciles: total treatment time matches the session that happened, and any timed-code minutes are consistent with it and with each other.
  • No borrowed observations: anything a patient, parent, or caregiver told me is written as their report, not as my clinical finding.
  • Details survived transcription: name, date, laterality, body part or target, diagnosis references, and pronouns are all correct.
  • Skilled service is visible: the note names what I did — the cue level and how I faded it, the grading decision, the safety judgment — not just what the patient did with “support” or “assistance.”
  • The reasoning sounds like a decision: at least one sentence says why I kept, changed, or cut something today, in language no bystander could have written.
  • The assessment is about today: it says something the previous note could not have said. If it would have been equally true last week, it is not an assessment yet.
  • Diffed against the last signed note: no sentence rode along unchanged from the previous session unless it is still precisely and independently true.
  • Deviations are in the record: anything I dropped, added, shortened, or modified mid-session appears, with the clinical reason.
  • Nothing important is missing: I checked for errors of omission — the refusal, the near-loss of balance, the caregiver training, the phone call — events the recording may have missed or the summary skipped.
  • The plan is a real intention: the plan line states what I actually intend to do next session, not a restatement of the long-term goal.
  • It survives the auditor read: one final pass reading as a stranger — this note alone, with no memory of the session, justifies skilled care today.

The workflow

Make the pass a habit, then make it smaller

A review pass that waits until nine at night inherits the exact problem AI drafting was supposed to solve — and it reviews against a memory that has already merged three sessions together. The checklist only works when it runs close to the session it checks. The second half of the discipline is feedback: the same tool making the same error in every draft is not a review problem, it is a configuration problem, and the review pass should be shrinking over time as you fix the source.

  1. 01

    Review at the point of memory

    Run the pass before the next session starts, or in the day’s documentation block at the latest. The accuracy items — invented numbers, missed deviations, borrowed observations — are only checkable while you still remember the session as itself, not as part of a blurred afternoon.

  2. 02

    Run the checklist in its fixed order

    Same sequence every note: numbers, time, attribution, transcription details, skilled language, assessment freshness, the diff, omissions, plan, auditor read. Order builds speed; improvisation rebuilds the slow re-reading habit the checklist exists to replace.

  3. 03

    Keep a two-week catch log

    Every edit you make before signing, tally it by failure mode. Two weeks of tallies tells you what your particular tool invents, softens, and copies — which turns a generic review pass into a targeted one, and gives you the evidence for the next step.

  4. 04

    Fix the tool, not just the note

    Recurring catches go upstream: adjust the tool’s template, instructions, or settings so the error stops being generated — many tools accept custom guidance like “never state accuracy percentages unless dictated” or “flag data not explicitly stated.” If the vendor offers no way to fix a recurring failure, that is product feedback worth sending, and worth remembering at renewal.

Who is responsible if an AI-generated note I signed contains an error?

You are. Medicare requires the author of a service to authenticate it, and CMS resolved the drafted-by-someone-else question for human scribes in 2017: reviewers look for the treating clinician’s signature, which affirms the note adequately documents the care provided. Signing an AI draft makes its contents your clinical record. That is exactly why the review pass is a clinical task rather than an administrative one — it is the step where you decide what you are willing to be the author of.

Do I have to disclose that a note was drafted with AI?

There is no single universal disclosure rule. Expectations vary by payer contract, state board, and employer policy, and consent for recording a session is a separate question from labeling the resulting note. What does not vary is authentication: whoever signs is the author, disclosed or not. Verify disclosure expectations with your payers, your board, and your compliance policy rather than assuming either answer — and if your practice adopts a position, put it in the documentation policy in writing.

Will using an AI scribe increase my audit risk?

The tool is not an audit category; documentation patterns are. The patterns payer review has always flagged — cloned entries worded like previous ones, objective data that does not reconcile with billed minutes, notes that describe unskilled activity — are precisely the ones an unreviewed AI drafting layer mass-produces. A disciplined review pass targets each of those directly, which is why the honest answer is: unreviewed, AI drafting concentrates risk; reviewed, it mostly relocates the time you were already spending.

How long should the review pass take?

Long enough to run the checklist honestly, which for a routine session note is short — most items are seconds once the sequence is habit. The truthful measure of an AI documentation workflow is total time from session end to signed note, review included, compared against what writing from scratch cost you. If review time is not falling after a few weeks, your catch log will show why, and the fix belongs upstream in the tool’s template or settings rather than in reading harder.

Can I let the AI write the assessment section?

It can draft one, but the assessment is where the review pass should assume the least and rewrite the most. The assessment is the section that exists to carry clinical judgment — why today’s performance means what it means, and what that implies for the plan. It is also where the two quietest failure modes concentrate: skilled reasoning softened into activity description, and yesterday’s conclusion copied forward. Treat the drafted assessment as a prompt for your judgment, not a substitute for it, and apply the test: does it say something the last note could not have said?

Primary sources

Bibliography / 8
  1. 01Complying with Medicare Signature Requirements (MLN905364)Centers for Medicare & Medicaid Services
  2. 02Medicare Program Integrity Manual Transmittal 713: Scribe Services Signature RequirementsCenters for Medicare & Medicaid Services
  3. 03Medicare Benefit Policy Manual, Chapter 15, §220.3: Documentation Requirements for Therapy ServicesCenters for Medicare & Medicaid Services
  4. 04Not All Recommended Fraud Safeguards Have Been Implemented in Hospital EHR Technology (OEI-01-11-00570)U.S. Department of Health and Human Services, Office of Inspector General
  5. 05Documentation: Cloned Documentation GuidanceNational Government Services (Medicare Administrative Contractor)
  6. 06Generative Artificial Intelligence (AI) for Clinicians in Audiology and Speech-Language PathologyAmerican Speech-Language-Hearing Association
  7. 07Practice Advisory: Emerging Technology — AI-Enabled Ambient Scribe TechnologyAmerican Physical Therapy Association
  8. 08Artificial Intelligence and Occupational Therapy: From Emerging Occupation to Educational, Practice, and Policy ImperativeAmerican Journal of Occupational Therapy (AOTA)

Written by Callie Editorial

Published August 3, 2026

Educational content, not legal, billing, or patient-specific clinical advice.