March 2026. A consumer electronics brand pulled its AI voice agent after two weeks. The reason? The bot was inventing product specs — not maliciously, not often, but enough to drive returns up 25% in under a month. Refunds, exchanges, furious customers on Google Reviews. Final tab: hundreds of thousands.
It's not an isolated story. A retail CX leader reported their agent was making up return policies and offering fictional discounts in 1.35% of tickets. For an SMB handling 2,000 calls a month, that's 27 problem calls. Every month.
And here's the kicker: most Quebec SMB owners we meet don't even know their agent can do this.
Why AI voice agents make stuff up
Let's be direct. A badly configured AI voice agent isn't "lying" — it's filling gaps. When it doesn't find an answer in its data, it generates the statistically most plausible response. Problem is: plausible doesn't mean true.
The numbers sting. According to early-2026 studies, 62% of businesses now cite hallucinations as the #1 barrier to AI deployment — more than fear of job losses (28%). Even the most advanced OpenAI and Google models still hallucinate between 15% and 20% on complex queries, according to data compiled by Ringly in its 2026 voice AI report.
Concrete translation for a Quebec SMB: if your agent handles 500 calls a month, 75 to 100 conversations could contain at least one fabricated piece of information. A phantom delivery policy. A price that's three revisions out of date. A promo that doesn't exist. An opening time that's wrong.
A real case we saw recently: a dental clinic in Laval had an agent that, on French-language calls, regularly suggested to new patients they could "pay in 3 interest-free installments." The clinic never offered that. Three weeks, 11 convinced patients. Disputed invoices. Complaints. Two 1-star Google reviews.
The real cost of a hallucination for a Quebec SMB
We run this math often with clients. It's brutal.
Take a restaurant processing 150 calls per week. At a 1% hallucination rate, that's 6 calls per month where the agent delivers bad info. If 2 of those involve a group reservation and the group shows up to a table that doesn't exist — the average loss (reputation hit + blocked tables + refunds) lands around $850 per incident.
Multiply by 12 months. That's $10,000 leaking out without anyone spotting the hole.
And we haven't even touched Quebec's Law 25 exposure. If your agent invents that "your data is never shared with third parties" while a processor sits outside Quebec — you just lied to your customers by proxy. And there, you leave the financial zone. You enter the legal one.
Before you sign with any vendor, the question isn't "is your agent smart." It's "what prevents it from inventing things." We laid out the right questions in our guide to the 10 questions to ask any AI voice agent vendor — read before committing.
Fix 1: Ground the agent in YOUR data (RAG)
RAG — Retrieval-Augmented Generation — is the technique that forces the agent to pull its answer from your knowledge base before speaking. No knowledge base hit? No invented answer. The agent says: "I can't find that information, would you like me to transfer you?"
According to Retell AI's research on hallucination mitigation, a properly implemented RAG reduces hallucinations by 80%. Not 20%. Eighty.
In practice, that means we connect the agent to:
- Your current menu (for a restaurant)
- Your up-to-date pricing grid (for a service)
- Your real hours — not last year's
- Your official policies, in Quebec French
The TECHMA team handles this plumbing. You don't configure the RAG yourself — we install it, sync it to your Google Sheet or booking system, and validate every key data point is wired in before launch.
Fix 2: Strict guardrails on off-limits topics
A well-designed agent has a list of "no-go zones." No medical advice without a professional. No commitment on a discount without approval. No legal answers. No final quote on a complex project.
When a customer asks something that touches a no-go zone, the agent escalates to a human or falls back to a fixed FAQ answer. It never improvises. Think of it like training a new employee: "when in doubt, transfer the call to Sophie." Except the agent does it every time, 24/7.
The ElevenLabs Conversational AI 2.0 platform now supports this kind of constraint at the conversation-node level — but you still have to configure the RIGHT guardrails for YOUR industry. A restaurant doesn't have the same no-go zones as a dental clinic.
Fix 3: Run simulation tests before go-live
This is the single most-skipped fix. And the dumbest one to skip, because it costs almost nothing.
Before an agent takes a single real customer call, we run 200 to 500 simulated conversations through it. Expected scenarios, weird ones, polite callers, confused ones, people who change their mind mid-sentence. At every turn we measure: did the agent give correct info? Did it hallucinate? Did it escalate at the right moment?
Agents that go through this step have a production error rate 4 to 6 times lower than agents launched cold. We've measured it across our own deployments.
Here's what we do with TECHMA clients: a simulation series in Quebec French AND English, including regional accents, local phrasing, and edge cases ("my mom booked for 6 but we'll be 8, and she paid with my boyfriend's card"). If the agent breaks there, we fix it BEFORE any real customer calls.
Fix 4: Continuous monitoring with divergence alerts
An agent that's working great today can drift tomorrow. Models get updated, data changes, edge cases emerge. The only way to know is to have eyes on it 24/7.
Good monitoring systems do three things:
- Flag conversations where the agent said something uncertain (low-confidence tone, ungrounded answer)
- Alert in real time if escalation or hang-up rates spike
- Generate a weekly report with the top 5 conversations to review
Without this, you discover problems three months in, when a customer emails to complain. Too late.
Fix 5: Human-in-the-loop for high-stakes cases
Some calls should never be 100% automated. Contracts over $5,000. Large refund requests. Medical or legal questions. Formal complaints.
The rule: the higher the stakes, the lower the confidence in a unilateral AI decision should be. The agent takes the request, qualifies it, and routes to a human who decides. No pride needed here — this is what separates a good deployment from a Twitter fiasco.
We sometimes get "if a human intervenes everywhere, what's the AI for?" The answer: the human only steps in on 5-15% of high-value calls. The remaining 85-95% (confirming a reservation, giving hours, qualifying a lead) runs fully automated. The human focuses on what actually matters.
Why Quebec SMBs are particularly exposed
Three reasons. First, Quebec French. Most base models are trained on French-from-France, or formal written French. When a customer says "j'aimerais ça savoir c'est quoi vos prix pour un nettoyage," the agent can misinterpret "nettoyage" (dental cleaning? house cleaning? car detail?).
Second, Law 25. Quebec's requirements on consent, data localization and processing transparency are among the strictest in North America. An agent that invents a privacy answer exposes your business.
Third, size. Large enterprises have dedicated QA teams on their agents. An 8-person SMB? Not a chance. So without a partner handling monitoring, drift slips under the radar.
Despite these challenges, the business case stays massive. Forrester measured a 3-year ROI between 331% and 391% on well-executed voice AI deployments — with payback under 6 months. "Well-executed" is the operative word.
We've debunked several common misunderstandings about these agents in our article on the 5 most persistent myths about AI voice agents in Quebec. Read it if you're still on the fence.
Bottom line
Hallucinations aren't going away. But they can be contained to under 1% of interactions with the right setup — that's what TECHMA clients in production for over 6 months are living.
What it takes: properly wired RAG, strict guardrails, simulation testing before go-live, continuous monitoring, and human escalation on stakes-heavy cases. None of these 5 layers configures in 10 minutes. But together, they're the difference between an agent that makes money for your SMB and one that bleeds it.
At TECHMA, we install these 5 layers by default. You don't do the config yourself — our team handles it, from data plumbing to production monitoring. If you want to know what a well-built AI voice agent could look like for your specific context, book a free audit with our team. We look at your call volumes, identify risk zones, and tell you straight if AI is worth it for your case.
No corporate fluff. Just numbers and a plan.
