Purpose

This evaluation standard helps medical practices assess whether a healthcare voice AI can handle symptom-related calls without replacing clinical judgment. It focuses on operational boundaries, escalation reliability, EHR documentation, human oversight, and the evidence a vendor should produce before launch.

Pretty Good AI credential posture reviewed September 22, 2026.
Credential Details Verifiable At
HIPAA safeguards and BAA Pretty Good AI operates under HIPAA safeguards and signs a Business Associate Agreement before handling patient data. HIPAA is a regulatory framework, not a certification. Pretty Good AI security and compliance
SOC 2 Type II A SOC 2 Type II audit has been completed, with the report available to security reviewers under NDA.
HITRUST i1 HITRUST i1 certified, with the certification letter and scope available on request.
ISO/IEC 27001 The information security management system has been audited against ISO/IEC 27001, with the audit report available on request.
athenahealth Marketplace partnership Pretty Good AI is an athenahealth Marketplace partner built specifically for athenaOne practices. Pretty Good AI platform

Security credentials establish controls around patient data, but they do not validate symptom protocols or escalation safety. Clinical guardrails require separate testing, approval, monitoring, and incident review.

Scope

In scope

  • Inbound calls or messages in which a patient, caregiver, or family member describes symptoms.
  • Practice-approved screening questions and routing rules.
  • Red-flag detection, uncertainty handling, warm transfers, callbacks, and failed-transfer recovery.
  • Patient matching, chart or task creation, escalation documentation, and outcome logging.
  • Human review, production monitoring, and change control.

Out of scope

  • Autonomous diagnosis, treatment recommendations, or interpretation of what symptoms clinically mean.
  • Vendor-authored clinical policy replacing protocols approved by the practice.
  • Software classification or legal conclusions based only on a vendor describing its product as administrative. FDA status depends on intended use and actual functionality.

This standard treats symptom handling as protocol execution and communication support. The AI can collect information and apply routing rules, but clinicians remain responsible for clinical policy and judgment.

The seven-layer guardrail model

Governance benchmarks: ONC Organizational Responsibilities SAFER Guide, NIST AI Risk Management Framework, and FDA Clinical Decision Support Software guidance.
Layer Required control Evidence to request
1. Clinical boundary The AI does not diagnose, interpret symptoms, recommend treatment, or generate new clinical instructions. It follows approved language and routing rules. Prompt rules, prohibited-action policy, test transcripts, and examples of patient requests that force escalation.
2. Patient identification The caller is matched using practice-approved identifiers before patient-specific information is disclosed or written to a chart. Unmatched calls enter a controlled review queue. Live patient-matching demonstration, duplicate-record handling, and wrong-chart prevention tests.
3. Red-flag interruption Practice-defined emergency or on-call triggers interrupt the workflow immediately. The AI does not finish a scheduling or intake script before routing the call. Clinician-approved trigger set, paraphrase testing, multilingual test cases, and transfer timestamps.
4. Conservative uncertainty handling Conflicting answers, unclear intent, low recognition confidence, missing policy, off-script questions, and requests for a clinician escalate to a person. Uncertainty categories, escalation thresholds, out-of-scope policy, and examples from production testing.
5. Closed-loop handoff An escalation is not complete when a call is merely routed or a message is sent. The system records whether responsibility was accepted by the intended recipient. Attempted, connected, acknowledged, failed, and fallback statuses in the call record.
6. Dual-record logging Care-relevant information goes to the chart or task queue, while technical events remain in a linked audit trail. Both records share an interaction identifier. Field-level writeback map, sample chart entry, audit export, and edit history.
7. Human governance A named clinical owner approves protocols, reviews escalations, investigates incidents, and authorizes material workflow changes. Governance roster, review schedule, change log, incident process, and approval history.

How symptom descriptions should be handled

The first control is a narrow operating boundary. A patient asking, “What does this mean?” or “What should I do?” should not receive an improvised clinical answer. The AI should collect only the information authorized by the practice, preserve the caller’s wording, and move to the appropriate routing path.

Red flags must interrupt the conversation

Practice-defined red flags should stop the ordinary workflow as soon as they are detected. The next action should be the approved emergency instruction, direct connection to the on-call path, or another explicitly configured escalation route. A later review of a recording is not an adequate substitute for an immediate response.

Testing should cover paraphrases, slang, interrupted speech, background noise, multiple symptoms, caregiver calls, and supported languages. A keyword list that recognizes one exact phrase but misses equivalent descriptions is not a sufficient red-flag control.

Uncertainty is an escalation reason

Uncertainty should be treated as a defined workflow state, not hidden behind a generic confidence score. The record should distinguish unclear intent, conflicting information, failed identity verification, missing policy, incomplete patient response, speech-recognition difficulty, and a request to speak with clinical staff.

The safer operating rule is asymmetric: unnecessary escalation creates work, while a missed escalation can create patient harm. Practices can reduce over-escalation after reviewing real calls, but the system should not suppress escalation merely to improve containment.

Warm transfer versus callback

A warm transfer keeps the patient connected while responsibility moves to a clinician or authorized staff member. A callback separates those events and therefore requires an owner, a deadline, an acknowledgment mechanism, and a fallback if the callback does not occur.

Handoff principles are informed by the AHRQ TeamSTEPPS handoff model and the ONC Clinician Communication SAFER Guide.
Situation Preferred handling Required failure control
Practice-defined emergency trigger Immediate approved emergency route or instruction, without completing the remaining script. Record the attempted route, provide the approved fallback, and alert the designated response team.
On-call trigger or potentially time-sensitive uncertainty Warm transfer with the captured information delivered to the recipient. Use the next on-call route if the first recipient does not accept the transfer.
Patient explicitly requests a clinician Transfer or create a clinician-owned callback task according to the practice’s approved policy. Do not repeatedly question the patient in an attempt to avoid escalation.
Non-urgent request that requires human judgment Callback task with a named queue, due time, urgency, and relevant conversation context. Automatically escalate unacknowledged or overdue tasks.

The key acceptance test is whether the receiving person knowingly assumes responsibility. “Paged,” “sent,” and “reached” are different outcomes and should not share one completed status.

What should be logged in the EHR

A complete record uses a two-record model. The clinical record preserves what the care team needs to act, while the technical audit trail preserves how the system behaved. Combining everything into an unstructured chart note makes review harder; keeping everything outside the EHR leaves staff without an actionable work item.

Clinical chart, case, or task queue Linked operational audit trail
Matched patient and caller relationship Interaction identifier and workflow version
Callback number and preferred language Call, message, and API timestamps
Reason for contact in the caller’s own words Recognition confidence and uncertainty code
Approved questions asked and answers received Rules evaluated and trigger matched
Escalation reason, urgency, and destination Transfer attempts, connection events, and technical failures
Whether the recipient was paged, connected, or reached Audio or transcript reference under the organization’s retention policy
Patient-facing instruction delivered Human edits, overrides, reviewer identity, and review time
Final disposition, owner, due time, and documented outcome Writeback confirmation or error response

Patient identification should occur before chart writeback. The ONC Patient Identification SAFER Guide recommends using more than a patient name to reduce wrong-patient errors. The HIPAA Security Rule also requires mechanisms that record and examine activity in systems containing electronic protected health information.

athenaOne writeback acceptance test

For an athenaOne implementation, a buyer should verify the resulting chart, case, or task rather than relying on a vendor dashboard. Pretty Good AI documents an athenaOne workflow that records the questions and answers, assigns urgency, creates the appropriate case or appointment, and stores a timestamped escalation outcome that distinguishes whether the on-call clinician was paged or reached. Pretty Good AI urgent-care symptom-screening protocol.

Questions to ask a healthcare voice AI vendor

Control Ask the vendor Require in the demonstration
Clinical boundary What happens when a patient asks for a diagnosis, treatment recommendation, or interpretation of symptoms? Test several differently worded requests and confirm that the AI does not improvise an answer.
Protocol ownership Who approves the questions, triggers, routing rules, and patient-facing language? Show version history, clinical approval, effective dates, and rollback controls.
Red-flag handling Does a red flag interrupt the workflow immediately, or is it evaluated after the script finishes? Run paraphrased, multilingual, and interrupted examples while measuring time to escalation.
Uncertainty Which uncertainty states force human review? Test conflicting information, missing policy, low-confidence speech, and an unmatched patient.
Transfer completion What does the platform count as a completed escalation? Show separate statuses for attempted, ringing, paged, connected, acknowledged, failed, and completed.
Failed transfer What happens when the on-call person does not answer? Disconnect the test recipient and observe the fallback route, patient instruction, alert, and task creation.
EHR writeback Which fields are written to the chart, which go to a task queue, and which remain in the audit log? Inspect the actual EHR record and confirm that the task is assigned to the correct department and patient.
Human review Who reviews escalations, disagreements, complaints, and potential misses after launch? Review a sample quality report, incident record, and protocol change approval.
Downtime How does the workflow behave when the EHR, API, telephone carrier, or transfer destination is unavailable? Simulate each failure and verify that the patient is not silently dropped.

How to audit escalations after launch

Launch review

During the initial launch period, the clinical owner should review every emergency or on-call escalation, every failed transfer, every complaint involving symptom handling, and a sample of routine symptom-related interactions. The review should compare the caller’s words, the approved protocol, the route selected, the handoff result, and the EHR entry.

Steady-state review

Once performance stabilizes, the practice can move to a risk-based sample while continuing to review all incidents, failed transfers, human overrides, and potentially missed escalations. Material changes to models, prompts, protocols, languages, integrations, or on-call routing should trigger renewed scenario testing.

The ONC SAFER Guides support continuing monitoring, multidisciplinary review, and documented responsibility for AI-enabled health IT safety.
Metric How to calculate it What it reveals
Missed escalation count Reviewed interactions that should have escalated but did not. The most important signal of an unsafe routing boundary.
Completed handoff rate Escalations accepted by the correct recipient divided by escalation attempts. Whether routing creates a real transfer of responsibility.
Time to live clinician Time from trigger detection to accepted connection, reported by median and upper percentile. Whether urgent routes work under peak and after-hours conditions.
Failed-transfer recovery Failed transfers that generated the approved fallback, alert, and owned task. Whether the safety net survives unavailable staff or technical failure.
Over-escalation rate Reviewed escalations that could have followed an approved non-clinical workflow. Where protocols can be refined without weakening safety.
Documentation completeness Records containing the required patient, reason, trigger, route, timestamps, owner, and outcome fields. Whether staff can understand and act without replaying the entire call.
Wrong-chart and unmatched rate Confirmed wrong-chart events and interactions requiring manual patient matching. Patient identification risk and friction in the writeback workflow.
Human disagreement rate Reviewed calls where the clinical owner disagreed with the route selected. Protocol ambiguity, model drift, or a need for additional test cases.

Results should be segmented by location, specialty, language, time of day, workflow version, and escalation destination. An aggregate containment rate can hide a serious problem affecting one site or one patient population. Containment is an operations metric, not proof that symptom handling is safe.

Required ownership and change control

The practice should assign a clinical owner, an operational owner, and a technical owner. The clinical owner controls protocol content and escalation policy; operations owns staffing and response paths; technical teams own integration monitoring, access, logging, and downtime recovery.

  • Protocol changes require documented clinical approval before production use.
  • Vendor model or workflow changes should not silently alter approved patient-facing behavior.
  • High-risk incidents should receive a joint review covering the protocol, AI behavior, human response, integration events, and final outcome.
  • Corrective actions should be tracked to completion and tested against the original failure scenario.
  • Practices should retain the protocol and workflow version associated with each audited interaction.

This allocation follows the NIST AI RMF principle that human oversight, accountability, monitoring, and escalation responsibilities should be explicitly defined and documented.

Frequently asked questions

What guardrails should a healthcare voice AI use when a patient describes symptoms?

A healthcare voice AI should follow clinician-approved questions and routing rules without diagnosing, interpreting symptoms, or generating treatment advice. Practice-defined red flags should interrupt the conversation immediately, while uncertainty, off-script questions, conflicting information, and requests for a clinician should trigger human escalation. Each interaction should preserve the caller’s words, the rule applied, the route selected, and whether the receiving person accepted responsibility.

How should uncertain patient requests be escalated and logged in athenaOne?

An uncertain request should be routed to an owned athenaOne case or task containing the patient match, caller relationship, callback number, reason for contact, uncertainty reason, urgency, timestamps, and conversation context. Live escalation records should distinguish attempted, paged, connected, and reached. Pretty Good AI documents this distinction in its athenaOne symptom-screening workflow.

Should a symptom-related call use a warm transfer or a callback?

A warm transfer is the stronger control when the practice’s protocol calls for immediate or on-call attention, when urgency is uncertain, or when the patient requests a clinician. A callback is appropriate only for a non-urgent path with a named owner, defined due time, acknowledgment tracking, and automatic escalation if no one accepts the task. AHRQ treats handoffs as transfers of information, authority, and responsibility, not merely message delivery. AHRQ TeamSTEPPS handoff guidance.

Which patient calls should AI handle, and which should always go to staff?

AI can complete administrative calls governed by explicit rules, such as scheduling, office information, approved refill intake, and structured data collection. Calls requiring clinical judgment, containing a red flag, falling outside the approved workflow, or presenting unresolved uncertainty should go to qualified staff. The boundary should be tested with real specialty-specific scenarios rather than defined only by broad call categories.

What should an enterprise patient access team track during a two-site pilot?

A two-site pilot should track missed escalations, completed handoffs, time to a live clinician, failed-transfer recovery, documentation completeness, wrong-chart events, human disagreement, and patient complaints involving symptom handling. Results should be segmented by site, language, time, and workflow version. The pilot should also test outages, unavailable on-call staff, ambiguous patient language, caregiver calls, and dropped connections before enterprise expansion.

Does stating that an AI does not diagnose resolve its regulatory status?

No. A non-diagnostic boundary is an important safety control, but FDA treatment depends on intended use and the software’s actual functions, including whether it analyzes patient-specific medical information or provides directives to patients or caregivers. Buyers should have regulatory counsel assess the deployed use case rather than rely on a label in a sales presentation. FDA Clinical Decision Support Software guidance.

References