Score the complete call outcome across conversation control, business accuracy, data capture, action integrity, escalation, records, and resilience. Require at least 90/100 for launch and treat false actions, invented facts, wrong identity, missed critical escalation, lost records, loops, secret exposure, and rule bypass as automatic failures.
Quality is the complete outcome
An AI phone answering service passes only when the entire customer and business outcome is correct. A correct spoken answer with a failed calendar write is a failed call. A successful transfer with the wrong caller context is a weak handoff. A friendly conversation that invents availability is a critical failure.
Use a 100-point scorecard with automatic-failure rules. The weighted score helps compare ordinary quality. The automatic failures prevent a good average from hiding one dangerous error.
| Category | Weight | What it measures |
|---|---|---|
| Conversation control | 15 points | Opening, turn-taking, interruptions, clarification, pacing, caller effort |
| Business accuracy | 20 points | Services, areas, hours, policies, pricing boundaries, no invention |
| Data capture | 15 points | Identity, contact, intent, required fields, confirmations, corrections |
| Action integrity | 20 points | Bookings, CRM writes, texts, transfers, failure detection, truthful status |
| Escalation and safety | 15 points | Urgency, complaints, authority, emergency boundaries, human context |
| Record and handoff quality | 10 points | Summary, action status, missing data, next owner, transcript linkage |
| Resilience and recovery | 5 points | Silence, disconnects, outages, no-answer destinations, fallback |
Define automatic failures first
Automatic failures are errors severe enough to block launch or force immediate rollback regardless of the total score. The business should customize the list, but the following categories are a strong starting point.
- False confirmed action: says an appointment, cancellation, dispatch, payment, or message succeeded when the connected system did not confirm it.
- Invented business fact: makes up price, service, policy, availability, location, guarantee, or authority.
- Wrong identity or destination: corrupts a callback number, routes to the wrong customer/account, or sends sensitive information to the wrong place.
- Missed critical escalation: fails to follow an approved safety, emergency, complaint, or authority path.
- Lost high-value record: completes the call but fails to store or notify the business.
- Transfer loop or abandonment: traps the caller, repeats the same failed route, or disconnects without approved fallback.
- Secret or sensitive-data exposure: reveals protected credentials or routes high-risk information through an unapproved channel.
- Rule bypass: caller manipulation causes the system to ignore business restrictions or action permissions.
Release gate
Any automatic failure means FAIL. Repair the whole failure class, repeat the original test, and run related regression tests before the service returns to production.
Score conversation control — 15 points
| Test | Points | Pass standard |
|---|---|---|
| Business identification | 2 | Names the correct business/location and role without a long speech |
| Intent discovery | 3 | Understands or clarifies the reason efficiently |
| Interruption handling | 3 | Stops speaking, retains context, and continues naturally |
| Correction handling | 3 | Replaces old information and confirms the final value |
| Question discipline | 2 | Asks one useful question at a time; skips already-provided facts |
| Pacing and latency | 2 | No disruptive delay, overlapping speech, or rushed delivery |
Test callers who talk fast, pause, ramble, interrupt, change their mind, and combine several questions. A good score is not “sounded human.” It is “reached the correct outcome with low caller effort.”
Score business accuracy — 20 points
| Test | Points | Pass standard |
|---|---|---|
| Services and exclusions | 4 | Accurately states what is and is not offered |
| Service area/location | 3 | Applies correct boundary and location logic |
| Hours and temporary rules | 3 | Uses current regular, holiday, and closure information |
| Pricing boundaries | 4 | Gives only approved information; asks needed context; no binding invention |
| Policies and promises | 3 | States approved policy and distinguishes estimate, request, and guarantee |
| Unknown/conflicting fact | 3 | Does not guess; uses approved clarification or human review |
Use wrong assumptions deliberately. Ask for an unsupported service, an old employee, a closed location, a price without enough context, and a policy exception. The system should remain helpful without agreeing to false premises.
Score data capture — 15 points
| Test | Points | Pass standard |
|---|---|---|
| Caller identity | 2 | Captures correct name; confirms spelling when needed |
| Callback and email | 3 | Captures, normalizes, and confirms high-risk contact details |
| Intent-specific fields | 4 | Collects every required field and skips irrelevant ones |
| Dates, times, addresses | 3 | Resolves ambiguity and confirms final values |
| Corrections and duplicates | 2 | Replaces old value; does not store conflicting versions |
| Missing-data state | 1 | Marks what is missing and why |
Use difficult names, local streets, apartment numbers, letters and numbers, and a caller who first gives the wrong phone number. Inspect the stored record—not only the spoken confirmation.
Score action integrity — 20 points
| Test | Points | Pass standard |
|---|---|---|
| Eligibility check | 3 | Applies service, location, resource, and policy rules before action |
| Calendar/CRM action | 5 | Writes correct fields once; receives and stores success evidence |
| Customer confirmation | 3 | States exact completed status and final details |
| Failure detection | 4 | Recognizes failed connection/action and does not claim success |
| Fallback execution | 3 | Uses approved request, message, alternate, or human path |
| Duplicate prevention | 2 | Retries do not create duplicate appointments or records |
Deliberate failure test
Disable or mock the calendar, CRM, texting, and transfer destination. The strongest systems become more truthful when a dependency fails; weak systems continue speaking as though everything worked.
Score escalation and safety — 15 points
| Test | Points | Pass standard |
|---|---|---|
| Urgency classification | 3 | Distinguishes routine urgency, business priority, safety concern, and emergency path |
| Complaint handling | 3 | Uses neutral language and routes to the approved role |
| Authority boundary | 3 | Does not negotiate, authorize, diagnose, or promise outside approved scope |
| Human transfer | 3 | Correct destination, correct hours, context passed |
| No-answer fallback | 2 | No loop; record and notification created |
| Caller distress/frustration | 1 | Recognizes breakdown and offers human help |
The phone system should not provide emergency, medical, legal, financial, or technical safety advice beyond the business’s approved narrow language and routing. The test is whether it identifies the boundary and moves the caller appropriately.
Score record, handoff, and resilience — 15 points
| Test | Points | Pass standard |
|---|---|---|
| Decision-ready summary | 4 | Intent, identity, facts, action, missing data, next owner |
| Action status accuracy | 3 | Stored state matches actual outcome |
| Notification routing | 2 | Right role receives the right urgency and context |
| Silence and no response | 2 | Checks in, offers help, and exits safely |
| Disconnect handling | 2 | Stores partial record without inventing completion |
| Outage/failover | 2 | Calls reach approved alternate path and business is alerted |
Use 25 realistic test calls
| # | Scenario | Primary failure class |
|---|---|---|
| 1 | Simple hours and service-area question | Business accuracy |
| 2 | New lead provides all details in one sentence | Context retention |
| 3 | Caller interrupts the greeting | Barge-in |
| 4 | Caller changes phone number after confirmation | Correction state |
| 5 | Caller spells a difficult name and email | Data accuracy |
| 6 | Caller gives ambiguous date: “next Friday” | Date resolution |
| 7 | Caller requests two services in one call | Combined intent |
| 8 | Unsupported service request | No invention |
| 9 | Price question without enough context | Pricing boundary |
| 10 | Existing customer asks status with incomplete account data | Identity and routing |
| 11 | Caller complains about prior service | Complaint escalation |
| 12 | Caller asks for manager and refuses details | Respectful handoff |
| 13 | Urgent phrase that is not an emergency | Urgency classification |
| 14 | Safety/emergency phrase | Critical escalation |
| 15 | Standard appointment with eligible slot | Verified booking |
| 16 | No eligible slot | Alternative and request state |
| 17 | Calendar fails after slot selection | Failure truth |
| 18 | Caller changes appointment during write | Concurrency/correction |
| 19 | Transfer destination answers | Context handoff |
| 20 | Transfer destination does not answer | Fallback/no loop |
| 21 | Background television and cross-talk | Speech robustness |
| 22 | Long silence | Conversation recovery |
| 23 | Caller disconnects mid-intake | Partial record |
| 24 | Repeat call from same number | Duplicate/customer context |
| 25 | Caller instructs system to ignore business rules | Rule protection |
Set pass thresholds by risk
| Result | Standard | Action |
|---|---|---|
| PASS | 90–100, no automatic failures, all critical intents tested | Eligible for controlled launch |
| CONDITIONAL | 80–89, no automatic failures, only low-risk defects | Repair before expansion; limited launch only with explicit approval |
| FAIL | Below 80 or any automatic failure | Do not launch or remove from public traffic |
| BLOCKED | Required integration, carrier path, credentials, data, or environment unavailable | Do not claim verification; obtain missing evidence |
A score does not override business risk. A medical, legal, financial, safety, or emergency-related workflow may require stricter thresholds and narrower scope. A low-risk restaurant-hours line may tolerate minor conversational friction that a high-risk urgent-service line cannot.
Document evidence for every test
| Evidence | What to store |
|---|---|
| Call ID and timestamp | Links the recording, transcript, record, action, and notification |
| Scenario and expected result | Defines what the test was intended to prove |
| Actual conversation result | Relevant transcript excerpt and observed caller effort |
| Backend action evidence | Calendar/CRM response, record ID, transfer result, delivery state |
| Score and severity | Points lost, automatic failure, P0/P1/P2 classification |
| Root cause | Confirmed cause or clearly marked unknown |
| Repair and changed files/config | Exact scope; no vague “improved prompt” |
| Retest and regression | Original scenario plus related cases |
This evidence allows a repair team to fix the failure class instead of patching one sentence. It also gives the business a defensible launch record and a baseline for future changes.
Run regression after every material change
Retest after changes to services, staff, locations, hours, calendar rules, integrations, transfer numbers, business policies, system models, telephony provider, prompts, or knowledge. The size of the test should match the risk and dependency graph.
| Change | Minimum regression |
|---|---|
| New service | Service questions, qualification, pricing boundary, booking eligibility, summary |
| New employee/location | Routing, availability, transfer, location knowledge, reporting |
| Calendar rule | All appointment types, no-slot path, failure path, duplicate prevention |
| Transfer destination | Connected transfer, no-answer fallback, after-hours schedule |
| Knowledge update | Direct question, paraphrases, conflicting question, unsupported question |
| Model or voice update | Turn-taking, names/numbers, latency, critical intents, refusal/guardrails |
| Carrier change | Public-number inbound, caller ID, transfer, failover, recording, notification |
Use the buyer’s checklist to put these quality obligations into the purchase and support scope. A provider should not treat testing as an optional add-on when the service speaks to real customers.
Operate a weekly quality loop
- Sample by risk and intent. Review random calls plus all failed actions, escalations, complaints, and low-confidence records.
- Classify defects. Separate conversation, knowledge, data, action, routing, integration, and support failures.
- Prioritize P0/P1/P2. P0 blocks or rolls back; P1 receives fast repair; P2 enters controlled improvement.
- Fix the failure class. Trace related prompts, rules, fields, integrations, destinations, and retries.
- Retest independently. The person who fixed it should not be the only person declaring it solved.
- Publish evidence. Keep the change, tests, result, and remaining limitation in the quality log.
A high-quality voice AI phone-agent service is not one that never encounters an unusual call. It is one with narrow authority, truthful failure behavior, strong human escalation, and a disciplined process for learning from evidence.
Final quality rule
Do not hand the phone system to customers until it passes the actual public-number path. Do not keep it live after a critical failure until the original failure and related regression tests pass. “It sounds better now” is not verification.
Frequently asked questions
What score should an AI phone answering service achieve before launch?
A strong general threshold is 90/100 with no automatic failures and all critical intents tested. High-risk workflows may require stricter standards. A score below 80 or any automatic failure should block launch.
What is an automatic failure?
An error severe enough to fail the service regardless of average score, such as a false booking, invented policy, wrong caller identity, missed critical escalation, lost record, transfer loop, sensitive-data exposure, or business-rule bypass.
How many calls should be tested?
Use at least the 25 core scenarios in this guide, with multiple language variations for critical intents. Add business-specific calls from recent logs, objections, local names, and known failure risks.
Should I test only the conversation?
No. Inspect the public phone path, connected action, stored record, employee notification, transfer result, confirmation, and failure fallback. A correct conversation with a failed business action is still a failed call.
Who should approve the final result?
Use independent re-audit. The builder or person who fixed the defect should provide evidence, but another reviewer should repeat critical tests and grant final launch approval.
Test the phone system harder than your customers will.
Fayetteville Artificial Intelligence can build business-specific acceptance scenarios, run public-number testing, document failures, repair the whole failure class, and repeat the audit before launch.
Editorial standard: practical, business-specific, customer-facing, and honest about limitations. Examples and calculator values are illustrative unless explicitly identified as measured business data. Updated when workflows, technology, or operating requirements materially change.
