Skip to main content
The Practice
Practice growthSeptember 26, 2026

Which AI Features in a Therapy EHR Are Worth Paying For

Every therapy EHR pricing page now has an AI line item, and every demo makes it look indispensable. This is a buyer’s framework for telling the difference: sort each feature by whether it removes a step from your week or adds a review step to it, run the trial tests that vendor demos are built to avoid, and use the federal transparency attributes as a free due-diligence question list.

Callie Editorial 17 min read
The AI line item
Net minutes

4 notes left

Close-the-day system

Capture

Objective data at point of care

Interpret

One clinical decision

Close

Sign, route, and clear exceptions

A finish line for every clinical day

At a glance

What you’ll leave with

  • The unit of evaluation for any AI feature is net minutes: the time the feature removes from your real week, minus the review time it adds. Generative features never take review to zero — APTA, ASHA, and AOTA all place verification of AI output squarely on the clinician — so a feature that drafts something is worth paying for only when the draft-plus-review loop beats doing the task yourself.
  • Sort every AI feature into one of three families before you look at price: features that remove a step you already do (easiest to justify), features that replace a step with a review step (worth it only where your hours actually go), and features that add a step you never did (default to no).
  • “AI-powered” is a marketing claim, not a specification — the FTC has said plainly that AI performance claims need substantiation like any other advertising claim. Test the claim yourself: run last week’s three most expensive tasks through the trial, time the full loop including review, and borrow the federal HTI-1 transparency attributes as your vendor question list.

Every practice management vendor now has an AI section on the pricing page, and every demo of it goes beautifully: the note writes itself, the schedule fills itself, the claim scrubs itself. The question facing a practice owner is no longer whether AI features exist or whether they can look impressive for four minutes — it is which of them, at add-on prices, actually return more time than they consume in your practice, on your caseload, with your payers. That question has a usable answer, and it does not require guessing which vendor’s model is smarter. It requires one test applied consistently: for each feature, what step does it remove from your week, and what review step does it add? Features that remove a step you already perform are usually worth evaluating. Features that replace a step with a review step are worth it only when the loop is genuinely faster than the task. Features that add a step you never performed are usually a demo, not a tool. This article builds that test into a feature-by-feature map, a set of trial tests a vendor demo cannot fake, and the transparency and contract questions to settle before the line item lands on your invoice.

The test

Judge every AI feature by the step it removes and the review it adds

Strip the branding off any AI feature and it does one of three things to your workflow. It removes a step you currently perform by hand — scanning a waitlist for a fit, checking a claim against payer edits before submission, retyping a dictated sentence into a field. It replaces a step with a different step — you no longer draft the note, but now you review a draft you did not write. Or it adds a step that never existed — a dashboard of “insights” someone now has to read, interpret, and decide whether to trust. Vendors present all three as time savings. Only the first is unconditionally a time saving. The second is a trade whose value depends entirely on whether the new step is faster than the old one for your caseload. The third is a cost wearing a feature’s name tag.

The reason the second family needs honest math is that for anything generative — drafted notes, drafted messages, summarized evaluations — the review step is not optional and not vestigial. APTA’s practice advisory on AI-enabled ambient scribes states that AI-generated documentation may contain errors of omission or addition and must always be reviewed by the clinician. ASHA’s guidance on generative AI frames these tools as support for clinical work rather than a substitute for clinical reasoning, and points clinicians to their ethical obligation to evaluate any technology they use. AOTA’s journal has pressed the same point for occupational therapy. Whatever the tool drafts, the clinician who signs owns — so the honest unit of evaluation is never “minutes the draft saves.” It is net minutes: the time the feature removes, minus the review time it genuinely requires, measured across a real week. A drafted note that saves twelve minutes of typing but needs six minutes of careful correction is a six-minute feature, and should be priced as one.

The centerpiece

The AI feature map: eight common features, judged the same way

The table below applies the test to the AI features that appear most often on therapy EHR and practice management pricing pages. It deliberately says nothing about any vendor’s implementation quality — that is what your trial measures. What it gives you is the right question to ask about each feature before the trial starts, so you evaluate the feature’s honest shape rather than its demo.

What each AI feature removes, what it adds, and when it earns its price

AI featureThe step it removesThe review step it addsWorth paying for when
Ambient scribe / AI-drafted session notesDrafting the note from memory after the session — often the largest after-hours block in a clinical weekA full clinical review of every draft before signing, plus the vendor diligence a listening tool triggers: BAA, training-data terms, recording consentDocumentation is your team’s biggest measured time sink, and clinicians commit to genuinely reviewing drafts rather than rubber-stamping them
Structured dictation and cleanupTyping — you narrate what happened and the tool formats it into the note’s fieldsA light proofread; the content is still yours, so the review is for transcription slips, not clinical substanceClinicians already think out loud well and the bottleneck is keystrokes, not deciding what to write
Billing code suggestionsLooking up codes and second-guessing timed-unit arithmetic on routine claimsVerifying each suggestion against the payer’s actual rules — the claim goes out under your name, not the model’sDenials or rework cluster on coding and unit errors, and the suggestions can be checked faster than the lookup they replace
Claim scrubbing and denial predictionManual pre-submission checks against known payer editsTriaging what the tool flags, including its false positivesYou submit enough volume that rework is a real weekly cost, and the flags prove more precise than your current checklist
Waitlist fill and schedule optimizationManually hunting through the waitlist when a cancellation opens a slotConfirming clinical fit, authorization limits, and family constraints before the offer goes outCancellations regularly go unfilled today and the suggestions respect the constraints that make a fill actually valid
Drafted patient messages and reply suggestionsWriting the first draft of routine, templated communicationReading every message before it sends, under the same PHI channel rules as any other patient communicationMessage volume is high and repetitive; a good static template library often captures most of the value for less money
Chart and evaluation summarizationAssembling history from a long chart before an evaluation, re-evaluation, or reportChecking each summarized statement against the source record — a summary that invents or omits is worse than reading the chartCharts are long, re-evals are frequent, and the tool cites where in the record each statement came from
“AI insights” analyticsUsually nothing — it adds commentary on top of reports you already haveReading, interpreting, and deciding whether to act on observations you did not ask forRarely as a paid add-on. A [KPI dashboard you defined](/resources/therapy-practice-kpis-dashboard) answers questions you actually have

Two patterns in that table are worth making explicit. First, the features with the cleanest case — dictation cleanup, claim scrubbing, waitlist fill — are mostly not generative glamour. They remove mechanical steps and add thin review, which is why they tend to survive the net-minutes test even in skeptical practices. Second, the features with the biggest headline promise — ambient scribes, chart summarization — are exactly the ones whose value swings hardest on your own behavior, because the review step is clinical, unskippable, and easy to underprice at demo time. Neither pattern says yes or no to any feature. It says: know which kind you are buying, and test accordingly.

The proof

Run last week through the trial, not the vendor’s sample data

Treat “AI-powered” as an advertising claim, because that is what it is. The Federal Trade Commission has told businesses in plain terms that claims about what an AI product can do need substantiation like any other performance claim, and that exaggerating or overpromising what the technology delivers is the kind of deception it pursues. You do not need to litigate any of that from a buyer’s chair — you just need the posture it implies: the burden of proof sits with the vendor’s product in your hands, not with the vendor’s demo in theirs. The way to collect that proof is the same bottleneck-first trial method that works for practice software generally, narrowed to the AI line items.

  1. 01

    Pick the three tasks the feature claims to attack, from your real week

    Not the vendor’s scenarios — yours. If you are evaluating an ambient scribe, that is your most common session type, your most complex one, and the one with the messiest room audio. If it is claim scrubbing, it is last month’s three most annoying denials. A feature evaluated on tasks you rarely do will look better than it is.

  2. 02

    Recreate each task in the trial with fictional or de-identified cases

    No real patient information enters any system before the business associate agreement covering your account is signed — trials included. Role-play the session, reconstruct the claim, rebuild the message thread. The realism you need is workflow realism, not real PHI.

  3. 03

    Time the full loop, including review and correction

    Start the clock when the task starts, and stop it when the output is genuinely done: note reviewed, corrected, and signed; claim verified against the payer rule; message read and sent. Then time the same task done your current way. The difference — net minutes — is the number the price has to justify.

  4. 04

    Read the errors, not just the average

    One invented data point in a drafted note matters more than ten formatting quirks, because it tells you the review step can never be skimmed. Note what kind of errors appear, whether they cluster somewhere predictable, and whether the tool lets you fix the cause in a template or setting rather than re-fixing the symptom every day.

  5. 05

    Price the winner at your volume, in writing

    Multiply net minutes by how often the task actually occurs, then set that against the add-on’s price at your clinician count — including what the price becomes at renewal and at twice your size. A feature that clears the bar for a six-clinician practice can fail it for a solo one, and the reverse.

Due diligence

The transparency questions — borrowed from federal certification

You are allowed to ask how the AI works, and there is now a federal template for what a straight answer looks like. ONC’s HTI-1 final rule created the first algorithm-transparency requirements for health IT certified under the federal certification program: systems with predictive decision support must let users see “source attributes” describing the intervention — up to 31 of them, covering things like the data used to develop it, how it was validated, its known limitations, and how its performance is monitored. Many therapy-specific EHRs are not ONC-certified and are not bound by any of it. That does not make the rule irrelevant to you — it makes it a free, federally drafted question list. A vendor whose product is certified can show you the attributes; a vendor whose product is not certified can still answer the questions, and how they handle being asked is itself due-diligence data.

Field checklist

08 items

What to ask about any AI feature before it touches your practice

  • What does this feature actually do, and what are its documented limitations — in writing, not in the demo narration?
  • Was it developed and validated on data from outpatient therapy settings, or is it a general-purpose model wearing a therapy interface?
  • Does our practice’s data — audio, notes, messages — train your models or your subcontractors’ models, and can we decline without losing the feature?
  • Is the AI feature covered by the business associate agreement, and which subcontractors (hosting, transcription, model providers) also touch our data?
  • Can we turn the feature off per clinician and per feature, so adoption is a choice rather than a migration?
  • How does the tool signal uncertainty or possible error, and how do we report the errors we catch?
  • What happens to AI-generated drafts and their source recordings at contract termination — are they part of our export, and what does deletion actually mean?
  • What is the price at renewal, at twice our clinician count, and if the feature moves into a higher tier?

The floor

Compliance items are gates, not scores

Nothing in the net-minutes math buys back a compliance failure, so run the gates before the scorecard. An AI vendor that receives protected health information on your behalf is a business associate under HIPAA, and HHS is unambiguous that a written business associate agreement must be in place before that information flows — there is no trial exception and no “we’re HIPAA compliant” exception. A feature that records session audio adds recording-consent questions that vary by state, covered in our consent walkthrough. And whatever any tool drafts, the clinician who signs is its author: the review obligations, and what payer reviewers actually punish in AI-shaped charts, are laid out in the sign-off article and the audit-risk article. If a vendor fails a gate, the feature’s brilliance is not a tiebreaker, because there is nothing to break: the answer is no.

Decide

A sane order for saying yes

Put the pieces in sequence and the decision stops being about AI at all. First, know where your hours go, because a feature can only pay for itself against a bottleneck you actually have. Second, run the gates — BAA, training-data terms, consent — and let a failed gate end the conversation. Third, trial the survivors against last week’s real tasks and keep the net-minutes numbers in writing. Fourth, buy the boring winners first: the step-removers with thin review usually fund themselves fastest, and they teach your team to trust or distrust the vendor’s AI before you bet your documentation workflow on the glamorous one. And revisit annually at renewal, because these features change faster than contracts do — this year’s demo-ware is sometimes next year’s genuine tool, and the reverse.

“An AI feature is not smart or dumb on your invoice. It is minutes in or minutes out, and only your week can tell you which.”

Which AI features in a therapy EHR are usually worth paying for first?

The unglamorous ones: features that remove a mechanical step and add only a thin review, such as structured dictation cleanup, claim scrubbing against payer edits, and waitlist fill suggestions. They produce positive net minutes with the least behavior change. Generative features like ambient scribes can be worth far more in absolute minutes, but only in practices where documentation is the measured bottleneck and clinicians hold the line on reviewing every draft.

Is an ambient AI scribe worth it for a solo therapy practice?

Sometimes — the math is just less forgiving. A solo clinician pays the whole add-on price against one caseload’s savings, so time the full generate-review-sign loop on your own sessions during a trial and compare it honestly against your current documentation time. If your notes already take three minutes because your templates fit your sessions, a scribe has little to remove. If documentation follows you home every night, it is the first feature to test seriously.

Do AI-drafted notes increase audit risk?

No payer rule prohibits AI-assisted documentation, and no reviewer sees which tool drafted a note. What payer review has always punished — and what an unmanaged drafting layer mass-produces — is uniform, cloned-looking notes and content nobody verified. The risk lives in the chart, not the tool, and it is controllable; our audit-risk article covers how reviews are actually selected and decided.

What should I ask a vendor about how their AI was built?

Borrow the federal template. ONC’s HTI-1 rule requires certified predictive decision support to expose source attributes covering the data it was developed with, how it was validated, its limitations, and how performance is monitored. Most therapy EHRs are not ONC-certified, but any vendor selling an AI feature can answer those same questions. Add the commercial ones: does our data train your models, which subcontractors touch it, and what happens to outputs when we leave.

Can I trial AI features with real patient data?

Not before a business associate agreement covering your account is signed. A vendor that receives protected health information on your behalf is a business associate under HIPAA, and the agreement must exist before PHI flows — trials included. Evaluate with fictional or role-played cases until it does; features that record audio also raise state recording-consent questions worth settling first.

How do I compare an AI add-on against just hiring help?

Convert both to the same unit. The add-on’s value is net minutes per week times the people affected; its cost is the subscription at your clinician count. Admin help is priced in hours and can absorb ambiguity and judgment calls that software cannot, but adds hiring, training, and PHI-access obligations of its own. For many practices the honest answer is a mix: software for the mechanical, repeatable steps, and people for the tasks where every case is a little different.

Primary sources

Bibliography / 10
  1. 01Keep Your AI Claims in Check (business guidance)Federal Trade Commission
  2. 02HTI-1 Final Rule (Health Data, Technology, and Interoperability)Office of the National Coordinator for Health Information Technology
  3. 03HTI-1 Decision Support Interventions (DSI) Fact SheetOffice of the National Coordinator for Health Information Technology
  4. 04Health Data, Technology, and Interoperability: Certification Program Updates, Algorithm Transparency, and Information Sharing (final rule)Federal Register
  5. 05Business Associates (HIPAA guidance)U.S. Department of Health and Human Services
  6. 06EHR Contracts Untangled: Selecting Wisely, Negotiating Terms, and Understanding the Fine PrintOffice of the National Coordinator for Health Information Technology
  7. 07Practice Advisory: Emerging Technology — AI-Enabled Ambient Scribe TechnologyAmerican Physical Therapy Association
  8. 08Generative Artificial Intelligence (AI) for Clinicians in Audiology and Speech-Language PathologyAmerican Speech-Language-Hearing Association
  9. 09Artificial Intelligence and Occupational Therapy: From Emerging Occupation to Educational, Practice, and Policy ImperativeAmerican Journal of Occupational Therapy (AOTA)
  10. 10Documentation Matters ToolkitCenters for Medicare & Medicaid Services

Written by Callie Editorial

Published September 26, 2026

Educational content, not legal, billing, or patient-specific clinical advice.