An AI receptionist test checklist should prove what happens to a real legal inquiry from the first ring through intake, booking, system writeback, and human escalation. A pleasant voice is useful, but it is not acceptance evidence. U.S. law firms need repeatable tests for accuracy, workflow completion, failure handling, governance, and total operating cost.

Law-firm evaluation playbook

Our recommendation: test TeleWizard against your hardest approved intake paths—not a staged vendor script.

TeleWizard is the strongest managed candidate when the firm wants complete intake, legal-system actions, 24/7 configured coverage, 50+ languages, firm-designated escalation, and an operating team that helps configure, test, launch, and improve the workflow.

Complete intakeFailure-path testingClio and LawmaticsHuman escalationAI Supervisor

Evaluate TeleWizard for Your Firm

AI receptionist test checklist connecting a law firm's calls, approved intake, booking, record updates, and human escalation
Conceptual TeleWizard editorial artwork showing legal professionals reviewing an AI reception test path; not product UI, a customer deployment, or evidence of results.

AI Receptionist Test Checklist: The Pass-or-Pause Verdict

A candidate passes only when it completes the agreed work, records the result correctly, and exposes any failure to a named owner. A fluent answer with a missing calendar event is a failure. A thorough intake written to the wrong contact is a failure. A transfer without caller context may be a partial pass, not a success.

Pass

The result matches the approved path, connected systems acknowledge the action, and the record is usable without avoidable rework.

Pass with condition

The path works, but a documented limitation, plan dependency, or controlled manual step remains acceptable to the firm.

Pause

The vendor needs a correction, clearer written scope, stronger evidence, or another test before production approval.

Stop

The system invents facts, crosses a legal boundary, hides failure, mishandles protected data, or cannot support a required path.

This framework is intentionally stricter than a demo. The U.S. AI call-center software buyer guide helps firms define the category and shortlist. This page supplies the evidence plan after a product reaches the shortlist.

Build the Acceptance Scorecard Before the Vendor Demo

Start with the firm’s current caller types, approved outcomes, connected systems, and human owners. Weight requirements before seeing a polished demonstration. Otherwise, voice quality and a long feature list can outweigh the unglamorous work that determines whether intake succeeds.

Use three scenario groups. Ordinary cases represent frequent calls. Boundary cases test uncertainty, possible conflicts, advice requests, unusual facts, and accessibility needs. Failure cases break calendars, transfers, integrations, audio, or destinations. Every required path needs an expected end state and evidence source.

Score the completed legal-intake result—not presentation polish
Area Required evidence Suggested weight Automatic stop
Scope and accuracy Approved answers, questions, uncertainty language, prohibited decisions, and faithful summaries 20% Invented fact, advice, representation, conflict, deadline, merit, value, or eligibility decision
Intake completion Required fields, valid unknown/declined states, confirmations, and usable next step 20% Required information silently omitted or changed
Connected actions Booking, transfer, contact, note, task, log, notification, acknowledgment, and deduplication 20% Silent action failure or incorrect protected record
Escalation Trigger, destination, context, fallback, owner, and after-hours behavior 15% No safe path for an advice request, emergency, or firm-required human review
Governance Access, recording, consent, retention, change approval, testing, incidents, and human oversight 15% Unreviewable change or unacceptable data handling
Economics Complete quote, usage model, implementation, support, optional services, and remaining staff work 10% Material cost or responsibility cannot be established in writing

The weights are a starting point, not a TeleWizard guarantee or universal standard. A criminal-defense firm may assign more weight to urgent escalation; a high-volume intake team may emphasize deduplication and writeback. Document the reason for each weight so the final decision is auditable.

Score evidence in two rounds. In the first round, the vendor may explain the intended configuration and identify what must be built. In the second, require the configured workflow to produce the actual result. Do not award full points for a roadmap item, a slide, or an action demonstrated in a different product. Record the environment, plan, permissions, scenario version, tester, timestamp, evidence link, result, limitation, and person responsible for correction.

Keep disqualifiers outside the weighted total. A high score should not compensate for one unacceptable legal-boundary or data-handling failure. Likewise, distinguish a product limitation from a configuration defect. A limitation may be acceptable when documented and covered by a safe operating path; a defect needs correction and a repeat test before approval.

Test 1: Can the AI Stay Inside the Firm-Approved Role?

Give the candidate a written role statement. It may greet callers, identify the contact type, collect approved information, answer approved administrative questions, schedule eligible consultations, route to designated people, and update configured systems. It must not decide conflicts, advise, establish representation, calculate deadlines, predict outcomes, judge merits, assign case value, or determine legal eligibility.

Then challenge the boundary. Ask the system whether the caller “has a case,” which chapter to file, whether a deadline has passed, what settlement is likely, whether the firm represents the caller, or what the caller should do immediately. A good result uses approved limitation language, preserves the request, and takes the correct administrative next step.

TeleWizard evidence to request: the approved knowledge and intake map, exact human-request and advice-request rules, firm-designated escalation destinations, and a demonstration showing that uncertain input is preserved rather than converted into a confident legal conclusion.

Test 2: Stress the Conversation, Not Just the Script

Real callers interrupt, backtrack, use speakerphone, spell names, give incomplete dates, ask several questions, decline to answer, and mix narrative with emotion. Use recordings or role-play that reflect actual audio conditions without exposing information the vendor is not authorized to receive.

Test accents, background noise, pauses, corrections, ambiguous names, multiple people speaking, rapid numbers, poor connections, repeated questions, and a caller who changes the purpose of the call. The system should confirm critical details, distinguish unknown from declined, avoid filling gaps, and recover without trapping the caller in a loop.

TeleWizard supports 50+ languages, but a count is not proof of your path. Phone language is typically selected for the call; enabled chat and messaging can usually switch. Test the firm’s terminology, names, numbers, dates, disclosures, booking, writeback, and escalation with qualified speakers.

Test 3: Can It Finish a Complete Approved Intake?

A message is not an intake. Choose one high-value, high-frequency path and define every required field, confirmation, permitted skip, disqualifying administrative branch, booking condition, transfer rule, and final record. For example, a personal-injury path may collect incident facts and treatment status as stated without deciding liability or case value. A family-law path may capture safe-contact instructions without offering strategy.

Run the path with complete information, missing information, a declined answer, an existing client, a wrong practice area, a possible conflict signal, and an urgent request. Inspect the transcript or summary and the structured record. Staff should be able to see exactly what the caller said, what remains incomplete, and who owns the next step.

TeleWizard’s managed model is a material advantage here. The implementation team helps discover, configure, connect, test, and refine the approved workflow. The firm still controls legal judgment and acceptance criteria.

Test 4: Prove Every Connected Action End to End

Do not accept a logo grid as integration proof. Ask the vendor to create or update the intended contact, write the approved note, log the call or intake, book the correct calendar, create the task, send the required notification, and show the acknowledgment. Then repeat the scenario with a duplicate contact, an unavailable calendar, revoked permission, invalid field value, and system outage.

TeleWizard publishes deep Clio and Lawmatics workflows and qualified MyCase and calendar connectivity. Clio’s official App Directory independently describes complete intake, contact updates, notes on contacts and matters, logging, scheduling, warm transfers with introductions, tasks, AI Supervisor, and managed onboarding. Exact actions still depend on the product, permissions, plan, and tested configuration.

Run the checklist against a real Clio workflow

Bring your intake fields, duplicate rules, calendars, task ownership, and failed-action expectations to the demonstration.

Inspect TeleWizard’s Clio Workflow

Test 5: Does Human Escalation Carry Context?

Define the triggers first: direct human request, advice request, possible conflict, complaint, uncertainty, sensitive exception, failed verification, integration failure, urgent firm rule, or caller type that must reach a person. For each trigger, define destination, hours, transfer type, context, fallback, and owner if nobody answers.

Test an available destination and an unavailable one. Verify whether the recipient receives the caller’s name, purpose, collected facts, urgency as stated, and the attempted action without disclosing more than the firm approved. TeleWizard routes to the firm’s designated people; it does not claim to supply a live receptionist on every call.

Use the AI receptionist human-escalation playbook for deeper context-transfer design.

Test 6: Carry the Same Rule Across Languages and Enabled Channels

If the firm enables phone, web chat, SMS, WhatsApp, email, Facebook, or Instagram, run the same intent through every purchased channel. Confirm identity and consent rules, disclosures, approved knowledge, required fields, booking, escalation, and record ownership. A channel is not complete because it can send a message.

Test whether context survives an approved channel change and how duplicate records are handled. Confirm optional services and usage in writing. TeleWizard can extend shared business logic across enabled channels, but availability and behavior depend on the configured service.

Test 7: Break the Workflow Deliberately

Reliable systems need safe failure states. Disable a calendar, remove a permission, send a malformed date, reject a CRM field, make the transfer destination unavailable, interrupt the connection, and provide an unsupported request. The desired response is not “never fails.” It is a visible, bounded failure with preserved context, a named owner, and an approved recovery path.

Adversarial and failure cases every finalist should run
Test Acceptable result Evidence to inspect
Calendar unavailable No invented booking; approved callback or alternate path with ownership Caller wording, alert, task, and later reconciliation
Duplicate contact Configured match or review path without uncontrolled duplicate creation Lookup, record chosen, fields changed, and audit trail
Transfer fails Approved fallback, context preserved, caller not falsely told the transfer succeeded Attempt status, destination, message or task, and owner
Ambiguous answer Clarification or unknown state rather than a fabricated fact Transcript, structured value, and summary
Out-of-scope request Boundary statement and approved administrative next step Classification, response, routing, and retained request
Immediate emergency Approved direction to 911 for U.S. police, fire, or ambulance emergencies Trigger rule and exact response; no claim of emergency dispatch

Test 8: Verify Coverage, Concurrency, and Peak Behavior

Simulate the firm’s ordinary and peak pattern across business hours, after-hours, no-answer, and overflow coverage. Measure answer time, conversation latency, completion, action success, transfers, and failures. Do not rely on an “unlimited” label. Capacity is subject to the written plan, carrier conditions, deployment constraints, destinations, and connected systems.

Repeat the load test after material changes to call flows, integrations, locations, languages, or campaigns. High volume magnifies both a good workflow and a bad one. A midsize firm should segment results by office, practice, language, coverage window, and outcome rather than accepting one blended average.

Test 9: Review Safeguards and Change Control

ABA Model Rule 1.18 addresses duties involving prospective clients, and Comment [4] advises limiting initial consultation to information reasonably necessary to decide whether to undertake the matter. ABA Formal Opinion 512 discusses competence, confidentiality, communication, candor, supervision, and fees when lawyers use generative AI. These are model-level authorities, not a vendor certification or universal law.

Review approved data, access, role disclosure, possible-conflict handling, advice requests, recording and consent, retention, deletion, vendor terms, incident response, accessibility, human oversight, and jurisdiction-specific obligations with qualified counsel. The NIST AI Risk Management Framework offers a voluntary structure for governing, mapping, measuring, and managing risk.

Ask who can change knowledge, questions, routing, destinations, permissions, and system actions; what review occurs; how the change is tested; and how rollback works. TeleWizard’s AI Supervisor can surface possible caller friction, objections, missed bookings, routing issues, and workflow gaps for human review. It is not a legal conclusion or compliance guarantee.

Test 10: Measure Outcomes and Preserve a Regression Set

Measure answer rate, completed intake, qualified booking under firm rules, show rate where attributable, transfer success, connected-action accuracy, after-hours completion, staff finishing work, and cost per completed intake. Define numerator, denominator, exclusions, time window, and source system before the pilot.

Keep representative ordinary, boundary, and failure calls as a regression set. Re-run them after changes to knowledge, questions, languages, calendars, permissions, destinations, or connected actions. Review a sample of real interactions by outcome—not only low-rated calls—so the team can detect silent writeback or routing problems.

Protect the test set itself. Use synthetic or properly authorized data for preproduction work, limit access, document retention, and avoid sending real prospective-client information into an environment the firm has not approved. When production validation is necessary, define the narrow sample, reviewers, deletion or retention rule, and incident path in advance.

Close the loop with change evidence. Each accepted change should name the problem, affected paths, approved adjustment, regression cases, reviewer, release date, and rollback. Re-measure the targeted outcome after release, but also inspect unaffected paths. This prevents a booking improvement in one practice from quietly damaging routing or disclosures elsewhere.

The law-firm intake KPI scorecard provides the formulas and attribution rules; this checklist focuses on whether the vendor can produce trustworthy evidence for them.

Price Only the Scope That Passed the AI Receptionist Test Checklist

TeleWizard costs less than hiring a dedicated full-time receptionist while delivering broader 24/7 coverage. Confirm that commitment with the firm’s custom quote and comparable configured reception and intake work. A full-time employee may perform in-office or judgment-heavy duties outside TeleWizard’s scope.

The U.S. Bureau of Labor Statistics’ May 2025 national data for receptionists and information clerks reports a mean wage of $18.97 per hour and $39,460 per year, with an $18.27 median hourly wage. Those are occupation-wide wages—not law-office total employer cost, staffing equivalence, software pricing, or a TeleWizard quote.

TeleWizard uses 3 credits per AI phone-agent minute, 5 credits per distinct in-call action per call, and 3 credits per call for enabled after-call work. Repeating the same in-call action during the same call adds no extra action charge. Carrier, recording, verification, memory, messaging, attachments, additional numbers, and other optional services may add credits.

TeleWizard pricing is tailored. Confirm included credits, hours, countries, languages, actions, after-call work, channels, integrations, implementation, support, AI Supervisor scope, and optional usage. Compare cost per accepted result plus remaining staff work—not a naked minute rate. Never invent a dollar-per-credit conversion.

A Practical 30-Day Test Sequence

Days 1–4

Map scope, callers, outcomes, prohibited decisions, systems, and owners.

Days 5–8

Freeze scenarios, evidence, weights, stop rules, and baseline definitions.

Days 9–13

Configure one complete path with approved knowledge and connected actions.

Days 14–18

Run ordinary, boundary, failure, language, and load cases.

Days 19–24

Correct issues, re-run the same set, and inspect every connected result.

Days 25–30

Review evidence, quote, responsibilities, launch conditions, and stop criteria.

This is a recommended evaluation sequence, not a promised TeleWizard implementation timeline. Scope, systems, permissions, legal review, data readiness, and acceptance testing determine the real schedule.

Finish with a written decision record. Preserve the final scorecard, accepted limitations, unresolved risks, quote assumptions, evidence locations, responsible owners, production acceptance criteria, review date, and conditions that would trigger suspension, rollback, or a new evaluation.

Test TeleWizard on the work your firm actually needs

Share one intake path, its systems, languages, coverage windows, failure cases, and human destinations for a tailored evaluation and quote.

Request a Tailored TeleWizard Scope

AI Receptionist Test Checklist: Frequently Asked Questions

How many scenarios should a law firm test?

Start with enough ordinary, boundary, and failure cases to cover every required outcome and stop rule. A smaller complete set with expected evidence is more useful than a large unstructured script library.

Should a firm test only the voice conversation?

No. Inspect the intake record, booking, transfer, task, notification, deduplication, failed-action path, and human ownership. A fluent call can still produce a bad operational result.

Can a vendor use its own demo script?

It can demonstrate the interface, but acceptance should use the firm’s approved scenarios and evidence. Every finalist should receive the same material test so the scores remain comparable.

Why is TeleWizard well suited to this process?

TeleWizard combines managed discovery, configuration, integrations, testing, launch, and refinement with complete approved intake, 24/7 configured coverage, 50+ languages, firm-designated escalation, and AI Supervisor insights.

Official Sources and Review Date

Product and pricing facts were reviewed August 21, 2026. Confirm the current written TeleWizard plan, connected actions, optional services, and test scope before launch.