The 'Audio-Only' AI Voice Agent Died on May 7, 2026: Why Quebec SMBs Will Lose to Multimodal Competitors Without Noticing | Agent IA Vocal
    Back to blog
    Opinion9 min readMay 19, 2026

    The 'Audio-Only' AI Voice Agent Died on May 7, 2026: Why Quebec SMBs Will Lose to Multimodal Competitors Without Noticing

    On May 7, 2026, ElevenLabs enabled sendMultimodalMessage. Why a multimodal AI voice agent is becoming the invisible edge for Quebec SMBs in 2026.

    MA

    Masdouk Adelakoun

    Cofondateur & CTO

    The 'Audio-Only' AI Voice Agent Died on May 7, 2026: Why Quebec SMBs Will Lose to Multimodal Competitors Without Noticing

    Introduction

    Picture the scene. It's 6:47 PM on a Wednesday in July. Mrs. Tremblay walks home from work, opens her basement door, and steps into 3 cm of water. She grabs her phone, googles "emergency plumber Longueuil," clicks the first result. On the site, a voice chat widget. She talks to the AI agent for 40 seconds, then sends a photo of the joint spraying water. The agent identifies the valve, confirms a technician can be there in 35 minutes, and books the slot. She closes the tab. The four other plumbers never even know she shopped around.

    This scene wasn't possible on May 1, 2026. It became possible on May 7. And most Quebec SMBs that already have an AI voice agent have no idea what just happened.

    What actually changed on May 7, 2026

    For those following the ElevenLabs changelog, the announcement looked innocuous. Version 1.1.1 of the JavaScript client exposed a new method called sendMultimodalMessage, and the conversational agents WebSocket server started accepting a new event type, multimodal_message. Three lines in the release notes. No marketing campaign.

    Here's what it means in practice. Before May 7, an AI voice agent on your website could only speak and listen. The visitor said something, the agent answered, you looped until the visitor clicked "Book appointment" or bailed. After May 7, the same agent can receive an image mid-conversation, look at it with a vision model, and adapt its spoken response in real time. Not a separate chat channel. Not an upload form. The visitor taps the camera icon in the voice widget and the image flows into the same conversation as the voice.

    This update is part of a broader wave. The same week, OpenAI announced GPT-Realtime-2 with GPT-5-class reasoning and a 128,000-token context window, while TechCrunch documented the arrival of a voice suite that blends speech, vision and reasoning. This isn't an experimental mode. It's the infrastructure that serious AI voice agent vendors are wiring up this week.

    The framing mistake that's going to cost Quebec SMBs

    Most SMB owners think of an AI voice agent as "a receptionist that never sleeps." Phone rings, agent answers, appointment booked, done. In that mental picture, multimodal is a developer toy. What's the point of an image during a phone call?

    The framing is wrong for a simple reason: the PSTN phone is no longer the majority channel. In 2026, visitors looking for a plumber, mechanic, vet or clinic land on the website first. On mobile, they open the widget. That widget can now be a multimodal AI voice agent. And when the visitor has a visual situation — water damage, a dashboard warning light, an injured animal, a mystery insect trap — the option to "send me a photo and I'll look at it while we talk" is the moment they decide to stay or leave.

    It's the classic mistake of the shopkeeper who thought the web would never replace the storefront. Storefronts still exist. But the first interaction happens somewhere else. The PSTN phone will still exist in 2030. But the first interaction of a customer in distress at 6:47 PM happens on the website widget.

    Voice and image converging into a single multimodal conversation thread — the May 7, 2026 ElevenLabs update visualized

    Voice and image converging into a single multimodal conversation thread — the May 7, 2026 ElevenLabs update visualized

    Why the Quebec market is uniquely exposed

    Three traits make Quebec SMBs particularly vulnerable to this shift.

    First, technical service sectors (plumbing, HVAC, electrical, auto mechanics, veterinary, property insurance) make up an outsized share of Quebec's SMB economy. These are precisely the sectors where a photo accelerates diagnosis. A customer who can send a picture of a valve, compressor, model nameplate or engine mechanism skips three qualifying questions and saves 4 to 7 minutes on the call. Multiply that by 50 calls a day.

    Second, Quebec SMB tech adoption has lagged Ontario and British Columbia for a decade. Fewer than 30% of Quebec SMBs with under 20 employees ran a chat widget on their site as of 2024. That lag flips into a competitive advantage for whoever moves first to multimodal: while competitors still ask the visitor to fill out a contact form, you're booking the slot within the same minute.

    Third, the language pressure. A French-speaking customer who senses they'll have to "get by in English" to explain a technical problem will quit. But if they can send a photo and explain in natural French — without having to translate "compression fitting" — the barrier disappears. Multimodal doesn't remove the need for a bilingual agent. It reduces the cognitive load of every interaction.

    The math of the "invisible loss"

    Here's the uncomfortable part. When an SMB loses a customer because of a missed call, they know it: the phone didn't ring, or it rang and no one picked up. It's measurable. Our May 2026 analysis pegged this cost at $126,000 per SMB per year in Quebec.

    But the multimodal loss is invisible. The visitor lands on your site, opens the widget, talks for 25 seconds, realizes they can't show the problem, leaves the tab, and clicks the next Google result. You'll never see this abandonment in your missed-call reports. It won't show up in your CRM. It won't show up in your voicemail stats. The only trace is a 25-second session in Google Analytics — anonymous, indistinguishable from a visitor who was "just browsing."

    My conservative estimates, based on conversational widget abandonment benchmarks documented by vendors since 2024, place that loss between 12% and 23% of possible web conversions for urgent technical services. On a monthly volume of 400 qualified visits, that's 48 to 92 customers silently picking the competitor who has the photo. At an average residential service ticket of $380, that's $18,000 to $35,000 per month evaporating. Per month. And nobody inside the company notices, because the phone, well, the phone kept ringing.

    "But our calls come in by phone, not on the web"

    The main objection I get when I lay this out to SMB owners is fair. "My customers call me on the phone. They don't shop on a widget." Three things to consider.

    One, your loyal customers call you. Your new customers — the ones deciding tomorrow whether to try your business for the first time — start on Google. The difference matters because that's exactly the cohort you want to convert.

    Two, the PSTN phone isn't going away. Multimodal complements it, it doesn't replace it. The multi-agent architecture we documented recently is precisely how you can have a phone voice agent and a multimodal voice agent on the web widget, sharing the same knowledge base and calendar, with zero configuration duplication.

    Three — and this is the most uncomfortable point — MMS over SMS and WhatsApp Business are already popular multimodal channels in Quebec. When your AI voice agent vendor doesn't support inbound attachments on those channels, you're leaving on the table a volume of requests that will never touch your main number.

    A Quebec plumber arriving at an emergency residential call he won through a multimodal voice widget interaction

    A Quebec plumber arriving at an emergency residential call he won through a multimodal voice widget interaction

    What to demand from your AI voice agent vendor in Q3 2026

    If you're shopping for an AI voice agent this quarter, or renegotiating with an existing vendor, here are the questions that will separate the serious from the amateurs.

    Question one: "Does your web agent support the multimodal_message event from the ElevenLabs SDK version 1.1.1 or higher?" If the vendor hesitates, asks to check, or answers "we're working on it for year-end," you have your answer. Vendors who own their stack were ready within 14 days of the announcement.

    Question two: "What's the added latency between the customer sending the image and the agent's adapted spoken response?" The realistic target is under 3 seconds using a fast vision model. Past 5 seconds, the experience feels clunky and the visitor slips back into doubt.

    Question three: "How is the submitted image stored, for how long, and under what Quebec Law 25 compliance policy?" A serious vendor has a precise answer. Multimodal multiplies the volume of personal data your voice agent processes — that's a topic to settle before deployment, not after. For a refresher on technical requirements, see our guide to the new ElevenLabs functions for Quebec SMBs in May 2026.

    Question four: "What honest comparison can you give me between the three main platforms on this capability?" Our Vapi vs ElevenLabs vs Retell AI comparison published in May 2026 already breaks down the multimodal support differences between the three platforms. A vendor who refuses that transparency isn't thinking about your interest.

    Conclusion: an uncomfortable prediction

    Here's my prediction for the next 18 months. By December 2027, the "audio-only" AI voice agent will be what a no-mobile-version website is today: technically still online, but read as a signal of lag by 80% of visitors. Quebec SMBs that move to multimodal before the end of 2026 will harvest 12 to 18 months of silent competitive advantage while their competitors keep counting missed calls instead of counting abandoned web sessions.

    May 7, 2026 wasn't a marketing event. It was a technical update that quietly shifted the boundary of what's possible. The first ones to notice aren't the ones who read a press release — they're the ones who saw, in their dashboard, their web conversion rate climb 15% without knowing why.

    If you want to explore what multimodal actually changes for your SMB — your website, your phone channel, your WhatsApp Business — let's take 20 minutes together. We handle the configuration and integration, you don't touch a line of code. That's our job, and it's included in the package.

    FAQ

    **Is multimodal available on classic phone calls (PSTN) too?** No, not directly. PSTN carries voice only. Multimodal is relevant on the web widget, mobile app, SMS/MMS and WhatsApp Business. The strength of a multi-agent architecture is to combine both worlds into one unified experience.

    **What's the added cost of adding multimodal to an existing agent?** On the platform side, image analysis typically adds $0.003 to $0.012 per image processed, depending on the vision model used. On a volume of 400 monthly images, that's $1.20 to $4.80 per month in platform cost. ROI, on a single urgent appointment recovered per month, is positive by several orders of magnitude.

    **Can my current website accommodate a multimodal voice widget without a rebuild?** In 95% of cases, yes. Integration consists of inserting a block of JavaScript in the footer. No backend, database or template changes are required. Our team handles the integration and testing on your environment.

    **How long does it take to put a multimodal agent into production?** For a standard SMB with an existing site, count 7 to 14 business days, including baseline information gathering, agent configuration, Law 25 compliance tests, and the validation phase. Zero code on your side.

    OpinionMultimodalElevenLabsQuebec SMB2026 TrendsVision AIAI Voice Agent
    Share