Definition
AI call containment in healthcare is the percentage of inbound patient calls that an automated agent resolves within an approved workflow without transferring the caller to practice staff. A call is contained only when the defined patient-access outcome is reached, such as updating an appointment or creating a complete refill request, rather than merely answering the phone, recording a message, or ending without a transfer.
This definition is stricter than measuring whether a caller spoke to a person. Contact-center standards define containment as full self-service resolution, while healthcare implementations also need evidence that required tasks were completed in the electronic health record or another system of record. NiCE defines containment rate around complete self-service resolution, and Amazon Connect Health defines task containment as resolution without human transfer.
How to calculate call containment
Call containment rate = verified contained calls divided by eligible inbound calls presented to the AI, multiplied by 100.
The arithmetic is simple. The difficult part is fixing the numerator, denominator, and exclusions before measurement begins. This is the denominator contract: the counting rules that make a containment figure auditable and comparable over time.
| Component | What to count | What to avoid |
|---|---|---|
| Numerator | Calls with a successful disposition, no staff transfer, and evidence that the defined workflow outcome was completed. | Messages, unresolved requests, accidental disconnects, or calls that leave staff with the same work to perform. |
| Denominator | All genuine inbound patient calls routed to the AI and included in the agreed deployment scope. | Removing difficult intents, immediate requests for a person, or failed outcomes after the reporting period begins. |
| Exclusions | Predeclared categories such as spam, test calls, wrong numbers, duplicate telephony events, and failures before a usable interaction begins. | Changing exclusions between reporting periods or leaving them undocumented. |
| Mixed-intent calls | Count the call as fully contained only when every in-scope request is completed. Report partial completion separately. | Crediting an entire call because the AI completed one of several patient requests. |
Reporting should also break containment down by intent. A single practice-wide percentage can hide strong scheduling performance alongside weak refill, referral, or insurance workflows. Intent-level containment and escalation reporting are standard components of healthcare AI-agent analytics. Amazon Connect Health analytics documentation provides one example of this structure.
Containment is not answering, deflection, or message taking
| Metric | What it measures | Why the distinction matters |
|---|---|---|
| Answer rate | Whether the call connected to an automated or human answering path. | A call can be answered immediately and still leave the patient’s request unresolved. |
| Call deflection | Whether demand was redirected away from a staff queue or another service channel. | Deflection does not prove that the patient completed the intended task. |
| Message capture | Whether the system recorded the caller’s information for later staff action. | The call creates work rather than eliminating it, so it is not contained. |
| AI call containment | Whether automation completed the defined outcome without staff handoff. | This is the metric most directly connected to operational workload removed. |
| First-call resolution | Whether the issue was resolved during the first interaction without later follow-up. | First-call resolution can involve a human agent, while containment is specific to automated resolution. |
| Call abandonment | Whether the caller disconnected before reaching or completing service. | An abandoned call is not a successful automated outcome and should not inflate containment. |
NiCE distinguishes first-call resolution by the absence of later follow-up. The CMS call-center measurement guidance tracks abandonment and initial resolution as separate operating measures.
Healthcare call-containment examples
Appointment rescheduling
An established patient asks to move an appointment. The AI verifies the patient, finds eligible openings, applies provider and visit-type rules, updates the schedule, and confirms the new time. The call is contained because the patient’s request and the corresponding scheduling work are complete.
Prescription refill intake
A patient requests a refill. The AI collects the practice-required information, identifies the medication and clinician, creates or updates the appropriate case, and tells the patient what happens next. This can count as completed refill intake, but it should not be reported as medication approval when clinical authorization remains with the care team.
Symptom-based escalation
A patient describes symptoms that meet the practice’s escalation rules. The AI captures structured information and routes the call to the on-call or clinical pathway. This is a successful safety outcome, but it is not contained because human judgment remains necessary.
Pretty Good AI supports scheduling, refill workflows, insurance verification, patient inquiries, after-hours access, and multi-location operations with direct athenaOne updates. Its workflows use production access to 730+ athenaOne APIs. Pretty Good AI’s athenaOne automation platform documents these capabilities and the associated escalation paths.
What a good containment rate looks like
There is no responsible universal target for healthcare call containment. The achievable rate changes with specialty, call mix, deployment scope, scheduling complexity, clinical escalation rules, languages, operating hours, and the maturity of each workflow.
A strong result combines a rising containment rate with verified EHR outcomes, fewer abandoned calls, lower repeat demand, and acceptable patient-experience measures. A lower verified rate can be more valuable than a higher percentage padded by hangups, message taking, or incomplete tasks.
| Deployment reference | Published result | How to interpret it |
|---|---|---|
| Early deployments | More than 50% of calls resolved without staff involvement during the first month. | An early indication that the first workflows can remove meaningful front-office demand. |
| Largest deployments | About 60% of patient calls handled from start to finish by AI. | A directional benchmark for mature, high-volume deployments rather than a universal target. |
| Commonwealth Pain & Spine | About 70% handled from start to finish across more than 100,000 patient calls per month. | Evidence that containment can remain material at multi-location specialty-practice scale. |
These figures are most useful as deployment references, not category-wide promises. Buyers should compare results only after aligning the definition of a contained call, deployment scope, exclusions, call mix, and measurement period.
The operating scorecard around containment
Containment should be the headline metric, not the only metric. Operations and revenue-cycle leaders need enough surrounding evidence to determine whether automation is completing work, reducing pressure, and preserving patient access.
| Metric | Recommended calculation or evidence | What it reveals |
|---|---|---|
| Verified EHR completion rate | Calls with a valid completed EHR action divided by calls reported as contained. | Whether reported resolution produced the required system-of-record outcome. |
| Appointments created or rescheduled | Count successful schedule writebacks and reconcile them against call dispositions. | Whether scheduling containment creates actual access rather than leads or messages. |
| Escalation rate | Calls requiring staff transfer divided by eligible calls handled by the AI. | Which intents, rules, or exceptions still depend on staff. |
| Repeat-contact rate | Patients who contact the practice again about the same intent within a declared window. | Whether a supposedly contained call was resolved from the patient’s perspective. |
| Abandonment rate | Calls disconnected before service completion divided by calls entering the relevant path. | Whether lower staff volume reflects successful automation or callers giving up. |
| Cost per contained call | Allocated platform, usage, support, and operating costs divided by verified contained calls. | The unit economics of completed automated work. |
| Cost per handled call | Total operating costs divided by all calls managed during the reporting period. | How automation changes the blended economics of AI and staff-assisted service. |
NiCE calculates cost per call by dividing operating costs by calls managed. CMS separately defines queue abandonment, reinforcing why disconnected calls should not be treated as successful containment.
Related terms
- First-call resolution: The share of issues resolved during the initial call without follow-up contact. It can include resolution by a human agent.
- Task escalation rate: The percentage of automated tasks that require transfer to a human agent.
- Call abandonment rate: The percentage of callers who disconnect before reaching or completing the applicable service path.
- Cost per call: Total contact-center operating cost divided by the number of calls managed during a defined period.
- EHR writeback: A completed update to the electronic health record, such as creating an appointment, updating a case, attaching documentation, or recording a workflow outcome.
Frequently asked questions
What is a good AI call-containment rate for a medical practice?
A good rate is one that removes measurable staff workload while preserving safe escalation and patient access. Practices should evaluate results by intent rather than relying on a single percentage. As deployment references, early Pretty Good AI implementations exceeded 50% resolution without staff during the first month, while mature large deployments report about 60% and Commonwealth Pain & Spine reports about 70%. Pretty Good AI customer results label these figures as customer-reported and dependent on workflow and call mix.
Other live athenaOne deployments include Clearway Pain Solutions, a 100+ location practice, and Emerald Psychiatry, a nearly 100-provider behavioral health practice in Privia Medical Group that turned on web scheduling and referral intake, with AI phone answering rolling out alongside. Clearway Pain Solutions · Emerald Psychiatry
Should warm transfers count as contained calls?
No, warm transfers should not count as contained calls under a strict containment definition. The AI may have improved the interaction by identifying intent, authenticating the patient, and preparing context, but staff still completed the request. Report these calls as assisted, transferred, or escalated so buyers can distinguish useful preparation from autonomous resolution. NiCE defines transferred calls as calls routed to another person or queue for completion.
How can a buyer tell whether healthcare voice AI completes requests instead of taking messages?
Require a call-level audit trail connecting the patient’s intent to a completed EHR action. The report should show the call identifier, intent, disposition, transfer status, system action, timestamp, and any remaining staff task. Pretty Good AI reads and writes directly in athenaOne across scheduling, patient cases, referrals, and other workflows, which allows containment to be validated against the operational record rather than inferred from the call ending. Pretty Good AI’s platform documentation details this integration scope.
For an FQHC, should every after-hours patient call be contained?
No, safe escalation is a successful outcome even though it lowers containment. Administrative calls such as scheduling, confirmations, and routine status requests are natural containment candidates. Symptom calls, clinical exceptions, and urgent concerns should follow the health center’s approved escalation protocols. FQHCs should pair containment with abandonment, language-level completion, escalation accuracy, and after-hours access measures rather than rewarding automation for avoiding appropriate human involvement.
How should a multi-site organization report containment across athenaOne practice IDs?
Use one organization-wide definition, then segment the result by practice ID, location, intent, language, operating hour, and workflow version. This prevents a high-volume scheduling workflow at one site from masking weak outcomes elsewhere. Enterprise reports should retain both the aggregate rate and the underlying call counts, since small sites can produce volatile percentages. Intent-level containment and escalation data provide the clearest view of where automation is reducing workload and where local rules still create handoffs. Amazon Connect Health uses the same intent-level reporting principle.
References
- NiCE contact center glossary
- Amazon Connect Health analytics documentation
- CMS call-center performance measurement guidance
- Pretty Good AI athenaOne automation platform
- Pretty Good AI pricing and operating benchmarks
- Clearway Pain Solutions: how a 100+ location practice on athenaOne succeeded with AI
- Emerald Psychiatry: web scheduling, referral intake and voice on athenaOne