How to Test an AI Voice Agent Before Signing: The 7-Call Protocol to Spot Mediocre Vendors (Quebec SMB Guide, May 2026) | Agent IA Vocal
    Back to blog
    Guide pratique10 min readMay 17, 2026

    How to Test an AI Voice Agent Before Signing: The 7-Call Protocol to Spot Mediocre Vendors (Quebec SMB Guide, May 2026)

    The 7-call protocol to test an AI voice agent before signing: Quebec accent, FR/EN switch, noise, Law 25. Quebec SMB guide, May 2026.

    MA

    Masdouk Adelakoun

    Cofondateur & CTO

    How to Test an AI Voice Agent Before Signing: The 7-Call Protocol to Spot Mediocre Vendors (Quebec SMB Guide, May 2026)

    Here’s the uncomfortable truth: a lot of AI voice demos are theater. Not all of them, but enough that you should assume the polished version you hear in a sales call may not be the one that answers your customers next month. In our experience, and based on what Quebec SMB buyers keep telling us, something like 60–80% of vendors can make a demo sound smooth, fast, and oddly charming — then the production agent arrives with more hesitation, weaker French comprehension, and awkward transfers. Yes, the “demo voice” problem is real.

    If you run a clinic, a law office, a garage, a property management company, or any SMB where missed calls turn into missed revenue, that gap matters. A bad AI receptionist does not just sound annoying. It creates friction, loses leads, and makes your business look disorganized. Worse, some vendors still push 12-month contracts around $6,000/year or more before you have tested the thing in conditions that resemble actual Quebec call traffic.

    So here is the practical fix: a simple 7-call protocol to help you test AI voice agent before buying SMB Quebec. No lab setup. No technical degree. Just seven calls that expose whether a vendor has a real production-ready system or a shiny demo built to survive exactly one scripted conversation. If you want a serious AI voice agent vendor evaluation, this is how you do it.

    Why demos fail as a buying tool

    A demo is controlled. The vendor usually knows the questions in advance, uses ideal audio conditions, picks the strongest voice, and avoids edge cases. That is normal in software sales — but voice AI is different because your customers do not behave like demo callers. They mumble. They switch from French to English. They interrupt. They call from a truck, a waiting room, or a Tim Hortons with a blender going full speed in the background.

    Modern systems have improved a lot, especially with newer real-time voice models such as GPT-Realtime-2. But even with stronger models and better speech stacks, execution still depends on setup quality, latency tuning, call routing, prompt design, fallback logic, and compliance choices. In other words: the model matters, but the implementation matters just as much.

    If you want more context on vendor screening, read the 4 questions to ask Canadian vendors. It pairs well with the protocol below.

    The 7-Call Protocol

    Run these seven calls yourself. Ideally, do them over 2–3 days, from different phones and environments. Do not warn the vendor in advance about the exact scenarios. If they only perform when prepared, that tells you something too.

    Call #1 — The Quebec accent test

    Use a strong regional Quebec accent. Not cartoonish. Just real. Speak at a normal pace and ask for something simple: opening hours, appointment availability, service area, pricing range. Then rephrase the same request in a more colloquial way. The goal is to see whether the agent understands natural Quebec French, not textbook Parisian French.

    What to listen for:

    • Does the agent correctly capture your intent on the first try?
    • Does it handle local pronunciation without forcing you to “speak robot”?
    • Does it answer naturally in French that sounds acceptable for Quebec callers?
    • Does it ask a clarifying question when unsure, instead of guessing wildly?

    PASS if: the agent understands the request with minimal friction, asks for clarification only when needed, and does not collapse when you use everyday Quebec phrasing.

    FAIL if: you have to repeat basic requests multiple times, simplify your accent unnaturally, or the agent keeps mapping your question to the wrong intent.

    This single test catches a surprising number of mediocre deployments. A vendor may claim bilingual support, but if the French side is weak in real Quebec conditions, your callers will notice immediately. And no, “it works better if callers speak more clearly” is not a serious answer.

    Call #2 — The FR↔EN switch test

    Start in French. Mid-conversation, switch to English naturally. Then switch back to French before the call ends. This is common in Quebec, especially in Montreal and in businesses serving mixed customer bases. Your AI receptionist should not behave like it has just been teleported into another dimension because you changed languages halfway through.

    What to listen for:

    • How quickly does the agent detect the language switch?
    • Does it continue the same task, or does it reset the whole conversation?
    • Does pronunciation remain clear in both languages?
    • Does the tone stay consistent, or does the agent suddenly sound like a different product?

    PASS if: the agent follows the switch smoothly, keeps context, and completes the task without forcing you to restart.

    FAIL if: it loses track of the request, answers in the wrong language repeatedly, or creates obvious confusion after the switch.

    This is one of the fastest ways to evaluate AI receptionist quality in Quebec. A lot of demos avoid this entirely because it exposes brittle language handling fast.

    Call #3 — The background noise test

    Make the call from a café, restaurant, lobby, sidewalk, or parked car with moderate ambient noise. You are not trying to sabotage the system with absurd chaos; you are simulating a real customer environment. Ask for a booking, a quote request, or a transfer to a department. Repeat one detail only if the agent asks — not because you feel bad for it.

    What to listen for:

    • Does speech recognition remain usable when noise is present?
    • Does the agent interrupt incorrectly because it mistakes noise for speech?
    • Can it capture names, dates, or phone numbers with reasonable reliability?
    • Does response speed degrade badly under messy audio?

    PASS if: the agent completes the task with only minor friction and handles moderate noise without repeated breakdowns.

    FAIL if: it mishears core details, talks over you constantly, or becomes practically unusable outside a silent office.

    Latency matters a lot here. If the vendor has not tuned turn-taking properly, noisy environments make the whole interaction feel clumsy. For more on that point, see the 800 ms latency rule.

    Call #4 — The interruption test

    Let the agent begin answering, then interrupt it mid-sentence with a correction or a new question. Do this politely but decisively. Real callers interrupt all the time, especially when they hear the answer is going in the wrong direction. A production-grade system should stop, listen, and recover without sounding offended — which, to be fair, would be impressive for software.

    What to listen for:

    • Does the agent stop speaking quickly when you cut in?
    • Does it process your interruption correctly?
    • Can it shift direction without replaying its previous script?
    • Does it maintain context after the interruption?

    PASS if: the agent yields the floor quickly, understands the correction, and continues the conversation naturally.

    FAIL if: it keeps talking over you, ignores the interruption, or restarts from a canned script as if nothing happened.

    This test reveals whether the underlying real-time stack is actually production-ready. Vendors love to say “low latency,” but you hear the truth when humans and the agent compete for the same half-second.

    Call #5 — The reasoning test

    Ask something slightly off-script. Not impossible, just not obviously preloaded. For example: “If I book today, what’s the fastest realistic timeline?” or “I need service for two locations; what would you suggest?” or “I’m not sure which option applies to me — what’s the best next step?” You want to see whether the agent can reason within boundaries, or whether it hallucinates with confidence.

    What to listen for:

    • Does the agent answer sensibly when the question is adjacent to the script?
    • When it does not know, does it say so clearly?
    • Does it offer a safe next step, such as taking details or transferring to a human?
    • Does it invent policies, prices, or availability that sound suspiciously made up?

    PASS if: the agent either gives a reasonable bounded answer or politely defers and escalates.

    FAIL if: it improvises false information, becomes incoherent, or traps the caller in circular responses.

    This is where the voice AI demo trap becomes obvious. In a demo, vendors often show only flows with predetermined intent. Real business calls are messier. Your AI agent does not need to know everything; it needs to know when not to pretend.

    Call #6 — The sensitive data test

    Give fake sensitive information deliberately — for example a fake SIN format or a fake credit card number format. Do not use real personal data. Then observe what the agent does. This is not just a technical test; it is a compliance and governance test, especially under Quebec’s privacy framework and Law 25 expectations. If the system casually invites, stores, or repeats sensitive data without proper controls, walk away.

    What to listen for:

    • Does the agent discourage collecting unnecessary sensitive data by voice?
    • Does it redirect to a safer channel when appropriate?
    • Does it avoid repeating sensitive numbers back out loud?
    • Can the vendor explain retention, storage, and access controls clearly?

    PASS if: the agent handles the situation cautiously, avoids unnecessary collection, and the vendor can document compliant practices.

    FAIL if: the agent accepts and repeats sensitive data freely, or the vendor gives vague answers about storage and privacy.

    For Quebec businesses, this is non-negotiable. Review the guidance and resources from the Commission d'accès à l'information du Québec and make sure the vendor’s approach is aligned with your obligations. A smooth voice is nice. A privacy complaint is less nice.

    Call #7 — The transfer test

    Ask for a human. Do it in a normal way: “I’d rather speak to someone,” or “Can you transfer me to the front desk?” Then test edge cases: ask outside business hours, ask when no one is available, or ask after the agent has already started collecting information. The handoff should feel fluid, not like being dropped into a void with hold music from 2004.

    What to listen for:

    • How quickly can the agent initiate the transfer?
    • Does it summarize context before handing off?
    • If no one is available, does it offer voicemail, callback, or message capture?
    • Does the caller need to repeat everything to the human?

    PASS if: the transfer is smooth, context is preserved where possible, and fallback options are clear.

    FAIL if: the handoff is slow, confusing, or forces the customer to start over from zero.

    This matters more than vendors admit. Even an excellent AI receptionist should know when a human is the better next step. The best systems reduce repetitive calls; they do not create a new layer of friction before a real person can help.

    How to score the 7 calls

    Keep it simple. Give each call a score:

    • 2 points = clear PASS
    • 1 point = mixed / acceptable with concern
    • 0 points = FAIL

    12–14 points: strong candidate. Worth discussing pilot terms.

    9–11 points: maybe, but only with a short pilot and very clear fixes.

    0–8 points: polished demo, weak production readiness. Move on.

    You can also weight the categories depending on your business. A bilingual clinic may care more about FR↔EN switching. A legal or financial office should weigh sensitive-data handling and transfers more heavily. A restaurant might care most about noise, interruptions, and speed.

    3 questions to ask AFTER the 7 calls

    1) Is the P95 latency SLA written into the contract?

    Do not settle for vague claims like “usually fast” or “near real time.” Ask for a contractual service level target around response latency, ideally at the P95 level, not just average performance. Averages hide bad moments. Customers remember bad moments. If the vendor refuses to define measurable latency expectations, that is a warning sign.

    2) Where is voice data stored, and who can access it?

    Ask where recordings, transcripts, and metadata are stored; how long they are retained; who has access; whether data is used for model training; and what controls exist for deletion and auditability. If the answer is fuzzy, legal review gets harder and operational risk goes up. This is especially relevant for Quebec SMBs dealing with customer information under Law 25.

    3) Can you run a 30-day pilot on a secondary line?

    This is the sanity test. If the vendor believes in the product, they should be open to a limited pilot on a secondary number or controlled call flow before a long commitment. Not every vendor will structure it the same way, but outright resistance usually means they know the live experience may not match the demo.

    If you want a deeper checklist of warning signs, read the red flags before signing.

    What a good vendor should say during evaluation

    A credible vendor will not promise perfection. They will explain limits clearly. They will tell you what the agent should handle, what should route to humans, how latency behaves under load, and what setup choices affect quality. They should also be transparent about the stack, whether they use providers like OpenAI real-time models, third-party telephony layers, or external TTS providers, and how they monitor performance after launch.

    They should also talk about implementation, not just software. At Agent IA Vocal, for example, installation is done by the TECHMA team for the client, because deployment details matter: call routing, fallback logic, business-hour rules, bilingual prompts, escalation paths, and testing in real conditions. Voice AI is not a logo plus a demo. It is operations.

    Common excuses mediocre vendors use

    • “That issue only happens in noisy environments.” — Right. Customers call from noisy environments.
    • “The bilingual mode is still being optimized.” — Then do not sell it as production-ready.
    • “Interruptions are hard for any AI.” — Less true in 2026 than it was two years ago.
    • “We can fix that after onboarding.” — Maybe. But you should not pay first to discover the basics do not work.
    • “Our best results are with scripted flows.” — Fine, but your callers are not scripts.

    A good AI voice agent vendor evaluation is not about catching tiny imperfections. It is about exposing whether the vendor has operational maturity or just presentation skills.

    Before you sign anything

    Do not sign a 12-month contract because the demo sounded smooth on Zoom. Test the live number. Test it in French. Test it in English. Test it with noise, interruptions, and an off-script question. Test what happens when the caller wants a human. And absolutely test how the system behaves around sensitive information. If a vendor cannot survive seven normal calls from a cautious buyer, imagine what happens with seven hundred real customer calls in a month.

    The safer path is simple: run a short pilot, on a secondary line, with measurable criteria. That gives you real data without risking your primary front desk experience. At Agent IA Vocal, that is exactly how we prefer to do it — a 30-day pilot on a secondary line so you can hear the system in your own business before making a longer commitment. No theater, no magic trick, just production reality.

    If you want to see whether an AI receptionist actually fits your business, book a discovery call with Agent IA Vocal. We will tell you where voice AI helps, where it does not, and how to test it properly before you spend a dollar more than necessary.

    AI voice agentQuebec SMBtestingprotocoldemoLaw 25May 2026
    Share