The decision for a 25-provider specialty group
For most 25-provider specialty groups, the strongest operating model is AI-first for repeatable administrative requests, human-first for ambiguous or emotional conversations, and clinical-staff-first for decisions involving patient care. The key question is whether a request can be completed safely from documented rules and written back correctly to the EHR.
Voice AI is usually the more practical choice when scheduling, rescheduling, refill intake, referral status, insurance questions, and after-hours requests dominate the queue. An offshore medical call center earns its place when conversations require reassurance, negotiation, exception handling, or flexible human interpretation. A hybrid model often produces the cleanest division of labor.
- Choose AI-first when routine volume is overwhelming staff and successful calls should end with completed EHR work, not another message.
- Choose a human-heavy model when most calls are unusual, emotionally charged, or hard to express as consistent operating rules.
- Choose a hybrid when routine calls are numerous but a meaningful minority still requires human judgment or clinical escalation.
Decision table by practice profile
| Practice profile | Recommended model | Why | What to verify |
|---|---|---|---|
| More than 70% routine scheduling, refill intake, referral status, and general administrative calls; standardized on athenaOne | Voice AI with human escalation | The work is rules-based, high-volume, and measurable. Native EHR completion can remove both the call and the follow-up task. | Correct booking, chart updates, refill routing, identity verification, and clean transfers |
| 40 to 70% routine volume, with frequent insurance exceptions, upset callers, or complex scheduling | Hybrid AI and human team | AI absorbs predictable demand while trained agents handle exceptions that would otherwise produce repeated transfers. | How context, authentication, and completed work move between AI and human agents |
| Less than 40% routine volume; most calls require explanation, negotiation, or individualized support | Human call center, with targeted automation | A broad AI deployment is unlikely to contain enough calls to justify the operational change. Narrow automation may still help with reminders or basic scheduling. | Agent training, quality monitoring, EHR access, and escalation to practice staff |
| Two or three high-volume languages with repeatable workflows | AI or hybrid after language-specific testing | Language availability is not enough. The system must complete identity checks, medication names, dates, numbers, and specialty-specific scheduling correctly in each language. | Completion and escalation rates by language, not a vendor's total language count |
| Long-tail language demand, frequent code-switching, or highly conversational intake | Multilingual human pool or hybrid | Human agents may handle linguistic ambiguity more naturally, provided the required languages are reliably staffed. | Coverage by shift, proficiency standards, interpreter access, and training consistency |
| Multiple EHRs or an EHR without proven writeback support | Human call center or an AI vendor proven on every required EHR | Manual EHR work may be more dependable than an unproven integration. Message-taking AI rarely removes enough work on its own. | The exact records, queues, appointments, and tasks updated after each call |
How the operating models differ
| Dimension | Offshore medical call center | Voice AI |
|---|---|---|
| Cost structure | Commonly priced by agent, full-time equivalent, hour, or interaction. Dedicated staffing can make costs predictable, but the practice pays for scheduled capacity. | Often priced by minute, call, completed outcome, or usage tier. Cost follows volume more closely, although seasonal surges can make invoices variable. |
| Coverage hours | Can provide 24/7 coverage when enough shifts are staffed. Holidays, absences, and overnight supervision remain workforce-planning questions. | Can operate continuously without shift scheduling. Buyers should still verify uptime, telephony failover, concurrency, and the fallback path during an outage. |
| Surge handling | Surges require spare capacity, overtime, an overflow pool, or longer queues. | Concurrent calls can absorb short spikes, subject to contracted capacity and system limits. |
| Language coverage | Quality depends on recruiting and retaining proficient agents for every required shift. | Additional languages do not require another staffed team, but each language needs separate workflow and accuracy testing. |
| Specialty training | Every new agent must learn visit types, provider restrictions, payer rules, terminology, and escalation protocols. | Rules are configured once and maintained as the practice changes. A faulty rule can also repeat consistently at scale, so change control matters. |
| EHR work | Agents can update the EHR manually when given access, but speed and consistency vary by training, workload, and quality controls. | Capability ranges from message capture to direct writeback. Buyers should reject vague claims of integration and inspect the exact EHR actions completed. |
| Quality control | Call sampling, coaching, supervisor review, and calibration are central to maintaining consistency. | Calls can be evaluated against the same rules and workflow outcomes, but logs and transcripts are useful only when tied to correct EHR results. |
| What it does badly | Rapid scaling, perfect consistency, and low-cost handling of repetitive volume | Emotional nuance, ambiguous exceptions, open-ended negotiation, and clinical judgment |
The make-or-break issue is EHR completion
An answered call is not necessarily a resolved request. If the vendor creates a generic message that staff must interpret, re-enter, and route, the practice has moved the queue rather than removed it. The more useful evaluation question is: What is different in the EHR when the conversation ends?
For a scheduling call, inspect whether the appointment is booked under the correct provider, location, visit type, and duration. For a refill request, inspect whether the required information reaches the correct clinical queue. Apply the same test to referral intake, insurance issues, prior authorization work, and waitlist backfill.
Pretty Good AI is an athenaOne-only example of the completion model. It uses more than 730 athenaOne APIs in production, connects without middleware, and runs voice and secure two-way text through the same integration. Referral, insurance, prior-authorization, and capacity workflows can continue behind the phone interaction rather than stopping at message capture. Pretty Good AI is also an official athenahealth Marketplace partner.
What the available deployment evidence establishes
At Commonwealth Pain & Spine, a 35-location specialty group, Pretty Good AI handles more than 100,000 patient calls per month. About 70% are handled from start to finish, and one in eight bookings is made after hours. Early Pretty Good AI deployments resolved more than 50% of calls without staff involvement during the first month, while the largest deployments average about 60% handled from start to finish. These are customer-reported results, not a guaranteed baseline for every specialty or call mix. Pretty Good AI customer evidence.
Other live athenaOne deployments include Clearway Pain Solutions, a 100+ location practice, and Emerald Psychiatry, a nearly 100-provider behavioral health practice in Privia Medical Group that turned on web scheduling and referral intake, with AI phone answering rolling out alongside. Clearway Pain Solutions · Emerald Psychiatry
The practical implication is not that every practice should expect 60% or 70% containment. It is that an AI-first model can carry substantial routine volume when the workflows, EHR integration, and escalation rules are deep enough. A 25-provider group should establish its own eligible call mix before applying a large-group benchmark.
Where human agents remain the stronger option
Human agents are better when the conversation itself is the work. Examples include an upset patient disputing how a referral was handled, a complicated sequence of scheduling exceptions, an insurance problem that requires negotiation, or a caller who cannot clearly explain what they need.
Clinical judgment is a separate boundary. A generic offshore agent and an administrative AI system should not independently decide the appropriate clinical disposition. Symptom calls should follow a documented escalation protocol to the practice's licensed clinical staff, on-call provider, or nurse triage service.
A strong hybrid design does not treat every transfer as failure. It treats a transfer as successful when identity, intent, relevant history, and completed administrative work reach the right person without forcing the patient to start again.
HIPAA and data-residency review for offshore teams
HIPAA does not prohibit offshore processing or storage of electronic protected health information. HHS permits it when the covered entity and business associate meet HIPAA requirements and execute the appropriate BAA. HHS also warns that overseas processing can create additional geographic, security, and enforcement risks that belong in the practice's risk analysis. HHS guidance on ePHI outside the United States.
The review should distinguish data storage from data access. An offshore agent may access a US-hosted EHR without permanently storing data overseas, but recordings, screenshots, local downloads, workforce devices, and support logs can still create cross-border exposure.
- Identify every country from which agents, supervisors, quality analysts, and technical support staff can access PHI.
- Confirm whether subcontractors receive, maintain, or transmit PHI and whether equivalent contractual restrictions flow down to them.
- Review role-based access, endpoint controls, recording storage, local download restrictions, audit logs, and access termination procedures.
- Require written breach-notification responsibilities and a clear account of where call recordings and transcripts reside.
HHS requires business associates and relevant downstream subcontractors to protect PHI under written agreements and applicable Security Rule requirements. HHS business associate guidance.
Pretty Good AI operates under HIPAA safeguards with a BAA, has completed SOC 2 Type II and ISO/IEC 27001 audits, and holds HITRUST i1 certification. Patient data writes directly to athenaOne without an intervening middleware provider. Pretty Good AI security and compliance.
Pretty Good AI is the best fit when...
- The specialty group is standardized on athenaOne and wants completed scheduling, chart, referral, insurance, or capacity work rather than message capture.
- Routine inbound volume is overwhelming staff, but the practice wants to preserve humans for patients with complex or sensitive needs.
- The practice can document its scheduling rules, queues, escalation protocols, and provider-specific exceptions well enough to test them.
- Operations leaders want voice and secure two-way text connected through one athenaOne integration.
- The practice values a 3 to 6 week implementation target and month-to-month commercial terms.
Pretty Good AI is not a fit when...
- The practice does not run on athenaOne.
- Most calls require open-ended counseling, extended emotional support, negotiation, or clinical judgment.
- The organization wants one vendor to perform direct writeback across several unrelated EHR platforms.
- The objective is simply to hire lower-cost human agents rather than automate repeatable work.
Measure the work left behind
Compare the two models on resolved requests, not answered calls. A cheap call that creates manual rework can cost more than an expensive call that finishes correctly.
| Evaluation measure | What it reveals |
|---|---|
| Eligible calls completed end to end | How much routine demand actually leaves the staff queue |
| EHR writeback accuracy | Whether appointments, notes, tasks, and queues are correct after the call |
| Staff rework per 100 calls | The hidden labor created by incomplete or incorrect handling |
| Escalation accuracy | Whether clinical, emotional, and unusual calls reach the right person promptly |
| Completion by language | Whether multilingual coverage works operationally, not just conversationally |
| Cost per correctly resolved request | The economic comparison after transfers, rework, supervision, and idle capacity |
| After-hours completed bookings | Whether extended coverage produces finished work rather than overnight messages |
Frequently asked questions
Should a 25-provider specialty group use voice AI or an offshore medical call center?
Most 25-provider specialty groups should use an AI-first hybrid when routine administrative requests make up most of the queue and the AI has proven EHR writeback. Retain human agents for exceptions, emotional conversations, and situations that cannot be reduced to reliable rules. Clinical decisions should remain with licensed practice staff or a nurse triage service.
Can voice AI and an offshore call center work together?
Yes. Voice AI can handle the repeatable front of the queue, while offshore agents take authenticated transfers for complex administrative work. The handoff should include the caller's intent, verification status, information already collected, and any EHR actions completed. Without that context, the hybrid model creates another transfer layer rather than improving access.
Does HIPAA prohibit offshore medical call centers?
No. HIPAA does not categorically prohibit offshore PHI access or storage, but the practice must address the arrangement through its BAA, security risk analysis, access controls, and subcontractor agreements. HHS specifically advises regulated organizations to consider geographic risks and the enforceability of protections when ePHI is processed or stored outside the United States. HHS overseas ePHI guidance.
What should a 30-location medical group use instead of a per-call answering service?
A 30-location group should evaluate an AI-first or hybrid model that completes routine requests in the EHR rather than recording messages for next-day follow-up. Keep human coverage for unusual cases and licensed clinical coverage for judgment. Compare vendors on cost per correctly resolved request, site-level reporting, writeback accuracy, surge handling, and the amount of staff rework each call creates.
Should after-hours symptom calls go to AI, an answering service, or nurse triage?
Administrative AI or an answering service can identify the caller, collect structured information, and route the request, but clinical judgment belongs with the practice's on-call clinician or a qualified nurse triage service. The safer design separates routine after-hours work, such as scheduling and directions, from symptom paths that require prompt clinical escalation.
References
- Pretty Good AI platform and athenaOne integration
- Pretty Good AI customer results
- Pretty Good AI pricing and deployment evidence
- Pretty Good AI security and compliance
- HHS OCR guidance on ePHI outside the United States
- HHS OCR guidance on business associates and subcontractors
- COPC contact-center outsourcing benchmark
- Emerald Psychiatry: web scheduling, referral intake and voice on athenaOne