Last Tuesday, a garage owner in Laval spent 40 minutes on the phone with a voice AI agent vendor. The conversation derailed on one question: "Which LLM do you want behind your agent? GPT-Realtime-2? Gemini 3.1? Claude? The new Qwen 35 that ElevenLabs just added?"
The owner hung up without signing. Not because the tech didn't interest him. Because someone was asking him to pick a brain for his agent without explaining what it actually changes for his shop.
He was right to be wary. On May 12, 2026, ElevenLabs added Gemini 3.1 Pro Preview, Qwen 35-35B-A3B, and Qwen 35-397B-A17B to its list of available LLMs for voice agents. That came on top of GPT-Realtime-2 (launched by OpenAI on May 7), Claude Sonnet 4.6, Gemini 3.1 Flash, GPT-4o-mini-realtime, and the Mistral family. Result: a Quebec SMB looking for a voice AI agent now has eight different brains to compare, each with its strengths, its bill, and its traps.
What this guide will save you from
We're not going to drown you in MMLU-Pro benchmark tables or "TTFT" times measured to the millisecond. Softcery engineers already published an 11-LLM comparison for English-speaking techies hosting their own H100 GPUs. That guide is useful, but it isn't about you.
You're a Quebec SMB owner. You handle between 200 and 2,000 calls per month. Your clients speak French, sometimes English. You're subject to Law 25 on personal information protection. And your phone automation budget hovers around 200 to 600 CAD per month.
So here's how to choose, based on the four things that actually matter for your shop: latency, monthly bill, French-English bilingual handling, and Law 25 compliance.
The eight brains in circulation (and which to skip right away)
Let's start with the full lineup as it shows up in the ElevenLabs console today:
- GPT-Realtime-2 (OpenAI, May 2026) — GPT-5-class reasoning, 128k token context, supports "let me check" style preambles. $32/M audio input tokens, $64/M output.
- GPT-4o-mini-realtime — OpenAI's budget option, very low latency.
- Gemini 3.1 Pro Preview (Google, added May 2026) — $2/M in, $12/M out, 124.9 output tokens per second. Excellent intelligence-to-price ratio.
- Gemini 3.1 Flash-Lite — Google's express version, $0.15/M in, ideal for high volume.
- Claude Sonnet 4.6 (Anthropic) — $3/M in, $15/M out, first-token latency around 1.24s.
- Qwen 35-35B-A3B (Alibaba, added May 2026) — mixture-of-experts model, sub-250ms streaming latency inherited from Qwen3-Omni Flash.
- Qwen 35-397B-A17B — the bigger sibling, better reasoning but slower.
- Mistral Large 2 — the European sovereignty pick, relevant only if your clientele is sensitive to non-US hosting.
Out of these eight, two can be ruled out immediately for most Quebec SMBs: Qwen 35-397B-A17B (the big one) because its voice latency blows past the 800ms rule that decides whether your clients hang up; and Mistral Large 2, which adds little if your data stays in Canada (your vendor can already host Claude or Gemini on AWS Canada Central).
That leaves six serious candidates. Let's see which brain fits which type of shop.
Restaurant doing 80 bookings/day: Gemini 3.1 Flash-Lite or GPT-4o-mini-realtime
A restaurant taking 80 reservations a day means roughly 2,400 calls per month. Half are wrapped up in under 90 seconds ("table for two, 7pm Friday"). The other half rarely exceed 3 minutes (changes, cancellations, menu questions).
That kind of call doesn't need deep reasoning. It needs speed, active listening, and the ability to understand "the table by the window, like last time" without grinding. For that, Gemini 3.1 Flash-Lite at $0.15/M input tokens does the job for around $90–130 CAD/month on the LLM side. GPT-4o-mini-realtime is in the same range, with a slight edge on interruption handling (the client who cuts the agent off to add "actually, make that four").
The trap here: picking Claude Sonnet 4.6 or GPT-Realtime-2 because someone sold you "the best quality." On 2,400 calls/month, the bill gap between Flash-Lite and Sonnet 4.6 can hit $400 CAD/month. For zero perceived benefit to your clients, who just want to book a table.
Clinic with urgent triage: GPT-Realtime-2 or Claude Sonnet 4.6
A dental or medical clinic is another game. When a patient calls saying "I've had this pain since last night, it's shooting up to my ear," your agent has three things to do in under 800ms: recognize urgency keywords, ask two triage questions (since when exactly? do you have a fever?), and decide whether to transfer, book an emergency slot, or schedule a callback from Dr.
That decision takes reasoning. Not "generate fluent text," light clinical reasoning. Two LLMs dominate here: GPT-Realtime-2 (GPT-5-class reasoning, supports preambles like "one moment, I'm checking Dr. Tremblay's schedule") and Claude Sonnet 4.6 (proven reputation for nuance and caution in medical contexts).
On 600 calls/month with a 3-minute average, the LLM bill lands around $180–280 CAD/month for GPT-Realtime-2, and $250–350 for Claude Sonnet 4.6. Pricier than the restaurant, but well-spent: a botched emergency triage that turns into a hospital visit is a patient lost for the next ten years.
Garage doing detailed estimates: Claude Sonnet 4.6 (no hesitation)
The garage owner who wants an agent capable of asking "constant noise or intermittent? when you accelerate or when you brake? what's the odometer reading?", then proposing a cost range and booking the lift for Tuesday morning — that owner needs a brain that holds a long conversational thread without losing context.
Claude Sonnet 4.6 is calibrated for this. Its context window, its ability to keep the thread after several detours, and the quality of its French Quebec paraphrasing ("you mean it knocks only when you turn left?") make it the default pick for consultative conversations.
The cost? On 400 calls/month at 5 minutes average, plan $220–320 CAD/month. If you want to dig into the math, we published the 5-step formula to calculate the ROI of a voice AI agent in Quebec.
Beauty salon with reminders and bilingual flow: Gemini 3.1 Pro Preview
Hair salon, esthetics, spa — clientele often 60-70% French-speaking, 30-40% English in Montreal, sometimes more complex in the suburbs. The agent must switch from French to English without grinding, handle reminders ("your appointment is tomorrow at 2pm, confirm or reschedule?"), and offer add-on services without sounding pushy.
Gemini 3.1 Pro Preview shines here. 124.9 output tokens per second, reasonable price ($2 and $12/M), and solid multilingual behavior thanks to its Google corpus training. According to Artificial Analysis benchmarks, it beats the median for reasoning models in its tier on speed.
The four numbers that should guide your decision (not the benchmarks)
Before signing with anyone, put these four questions to your vendor. Without good answers here, the LLM choice doesn't matter anymore.
1. End-to-end latency under 800ms? That's the threshold past which 40% of clients perceive lag and hang up. Ask for a recorded demo with timestamps.
2. Is the LLM hosted on AWS Canada Central or Azure Canada East? If the answer is "no, it's in the US," Law 25 compliance gets complicated. The vendor must justify any third-country data transfer.
3. What's the LLM bill alone on your projected monthly volume? Classic trap: a vendor quotes "$0.15/minute" without specifying what's included. Ask for the breakdown: STT, LLM, TTS, SIP transport, hosting, support. The LLM typically represents 30 to 50% of total cost.
4. What happens if OpenAI or Anthropic changes pricing or deprecates the model? In May 2026 alone, OpenAI announced GPT-Realtime-2 and deprecated the old Realtime API Beta. If your agent is locked into one model, you're exposed. A good vendor can swap LLMs in minutes without breaking your flow.
The three traps that cost Quebec SMBs 6 to 12 months
First trap: picking the "best" LLM instead of the right one. Claude Opus 4.7 tops the LMArena leaderboard, but at $15/M input tokens it's financially absurd for an agent taking restaurant reservations.
Second trap: getting sold an agent "powered by GPT-5" with no ability to swap the brain. That's the equivalent of buying a car with the hood welded shut.
Third trap: forgetting that integration isn't an isolated technical act. Wiring an LLM into your booking software, your CRM, your client records — that's where 70% of the work hides. At TECHMA, our team handles full integration on the client's behalf; never self-service. That's what separates an agent that "talks well" from an agent that "works well."
Our recommendation by profile, in one sentence
If 80% of your calls wrap in under 3 minutes and your volume exceeds 1,500 calls/month: Gemini 3.1 Flash-Lite.
If your calls need triage, reasoning, or medical/legal nuance: GPT-Realtime-2 or Claude Sonnet 4.6.
If your calls are long, consultative, and need to hold the thread for 5+ minutes: Claude Sonnet 4.6.
If your clientele is evenly bilingual FR/EN: Gemini 3.1 Pro Preview.
If you want to test something new and your vendor supports it: Qwen 35-35B-A3B, with the impressive streaming latency inherited from the Qwen3-Omni family.
Now what?
The LLM choice is just one of the 12 decisions to make when deploying a voice AI agent. Latency target, FR/EN architecture, CRM integration, escalation-to-human hours, fallback scripts, Law 25 compliance — each deserves its own conversation.
At Agent IA Vocal, we walk through this analysis with you on a 30-minute call. Not to sell you, to give you an honest opinion on which brain would fit your shop. If we're not the right fit, we tell you that too.
Because in the end, the right LLM isn't the one with the most parameters or topping the leaderboard. It's the one that picks up on the first ring, understands your client on the first try, and ends the conversation with a confirmed appointment in your calendar.
