Your AI Voice Agent Can Now See: Why April 2026's Multimodal Release Is Changing the Phone for Quebec SMBs | Agent IA Vocal
    Back to blog
    Tendances & Général7 min readMay 1, 2026

    Your AI Voice Agent Can Now See: Why April 2026's Multimodal Release Is Changing the Phone for Quebec SMBs

    April 2026: ElevenLabs added multimodal to AI voice agents. Here's what it actually changes for Quebec plumbers, brokers, and garages — beyond the hype.

    MA

    Masdouk Adelakoun

    Cofondateur & CTO

    Your AI Voice Agent Can Now See: Why April 2026's Multimodal Release Is Changing the Phone for Quebec SMBs

    "Send me a picture of the leaking pipe — I'll give you a quote right now." A plumber in Trois-Rivières said this to a customer on April 19, 2026. Except it wasn't the plumber speaking. It was his AI voice agent. And the photo? The agent actually looked at it, identified a cracked PEX fitting, then booked an emergency slot for the next morning — while the human plumber was finishing dinner.

    Welcome to the Quebec SMB phone in 2026.

    What just happened in April 2026 (and why nobody noticed)

    On April 1, 2026, ElevenLabs shipped what looked like a routine line in their changelog: "Multimodal message support added to Conversational AI." Five words, buried in a dev note. It's probably the most underrated change of the year for business phones.

    In plain English: an AI voice agent can now receive images during a call. Not afterwards. Mid-conversation. The customer sends a photo by SMS or a web form, the agent looks at it, understands it, and replies aloud while taking what it sees into account. The official ElevenLabs update also added guardrail events and multi-agent conversation tracking — but it's the multimodal piece that actually changes the economics of an inbound call.

    And it isn't an isolated move. OpenAI's GPT-Realtime added image input and SIP calling the same week. Three major providers, one month, the same direction. Voice is no longer blind.

    What does "seeing" on the phone actually look like?

    We asked roughly a dozen Quebec SMB managers in April. Nobody knew the feature existed. When we explained it, the reaction was always the same: "OK, but how do I actually use it?"

    Here are the four scenarios already running in production for clients in Montreal, Sherbrooke, and Quebec City:

    1. The customer sends a photo, the agent does a pre-diagnosis. Plumbing, mechanics, appliances, electronics: 60% of emergency calls start with "I don't really know how to describe it, it's just broken." With a photo, the agent recognizes the brand, the model, the type of failure, and dispatches the right technician with the right parts in the truck.

    2. The customer shows a document, the agent reads it. An insurance policyholder calling about a claim can now text a photo of their contract or damaged credit card. The agent extracts the right numbers, validates coverage, opens the file. No more spelling 16 digits over the phone.

    3. The customer shows an object, the agent confirms the part. Hardware stores, garages, auto-parts shops: a customer sends a photo of a broken component, the agent identifies the SKU, checks stock, reserves the item. No more "bring it in and we'll see."

    4. The customer shares their screen, the agent guides. IT support, online configurators, accountants buried in paperwork: a screenshot comes in, the agent sees the error message, gives the exact next step. Big U.S. contact centers have been testing this since last fall; it's now landing on small-business phones at a tenth of the cost.

    The four industries where the math actually changes

    Multimodal isn't relevant for every call. A restaurant booking is still simple voice. But in four sectors in Quebec, it redefines the unit economics of an inbound call.

    Home services. A plumber rolling a truck for nothing burns $80 to $150 in margin. With a pre-call photo, the rate of correct first-visit diagnoses climbs from roughly 55% to over 85% — that's what we're seeing with the clients we've migrated to multimodal in April. For deeper context on the economics of these trades, we've published a full breakdown of the 62% lost-call problem at plumbing and electrical shops.

    Insurance brokers. A file handled by voice averages 14 minutes. Add a photo of the policy, the licence, or the damage: 6 minutes. Fifty percent of capacity unlocked, no new hire. Recovering the 43% of calls falling into the void finally becomes a clean equation.

    Garages and body shops. Phone estimates were always fiction. Customer described a noise, mechanic guessed. Today, the customer sends a photo of the damage or a video of the overheating engine, and the agent classifies urgency before booking. The gap between the phone quote and the final invoice shrinks.

    Specialty retail. Hardware, supplies, parts: the killer line was always "I don't know what it's called." A photo solves that in two seconds. Average basket size goes up because the agent can suggest compatible accessories it can see.

    Un client envoie une photo a son agent IA vocal

    Un client envoie une photo a son agent IA vocal

    The Quebec context: Law 25, bilingualism, and Canadian-dollar invoices

    Three things to understand before signing for multimodal in Quebec.

    First, Law 25. A photo is personal data, especially when it includes a face, a licence plate, or a signature. Sending it via SMS or web form has to go through a locally hosted channel — or at minimum a vendor with a documented Canadian hosting agreement. Most current multimodal models still process images in U.S. data centres. That's not illegal, but it does trigger a privacy impact assessment. For sensitive sectors, we documented the on-premise option launched in April 2026, which solves exactly this problem.

    Second, bilingualism. An agent that sees also has to talk to your customer. A photo of a pipe captioned in Montréal-Nord franglais is still a real challenge. Both ElevenLabs and GPT-Realtime now handle two languages on the same image, but we recommend testing real code-switching before going live.

    Third, the bill. A second of voice costs roughly $0.002 USD. An image processed by a multimodal model? Between $0.015 and $0.04 depending on resolution and model. Tiny per call — but at 50 photos a day for a year, that's $700 stacking up. Worth budgeting upfront.

    What multimodal does not do (and please remember it)

    You'll read everywhere that multimodal voice is "going to change everything." That phrase has been recycled every quarter since 2023.

    Multimodal does not replace human diagnosis. A plumber on-site sees moisture inside the wall; the agent only sees what the customer aims at it. It also doesn't read emotions on a face in real time during a call — live video stream is still experimental, not production-ready. And it doesn't decide for you: an agent that "sees" a leak still doesn't dispatch a tech without human confirmation from the dispatcher.

    The right framing: multimodal removes wasted steps. It doesn't remove the trade. If anyone is selling you the opposite, ask whether they've ever run a body shop on a Saturday morning.

    Plombier prepare son intervention apres pre-diagnostic photo

    Plombier prepare son intervention apres pre-diagnostic photo

    How much it costs, and how to start (without getting fleeced)

    Three questions to ask before paying for multimodal:

    Do my calls actually need it? If 90% of your volume is appointment booking at a hair salon, multimodal won't change much. Although — for a photo of the haircut the customer wants to replicate, it's pure gold.

    Does my current voice agent support it? Not all of them. Check that your vendor integrates ElevenLabs Conversational AI v2026.04+, GPT-Realtime image input, or an equivalent model. If the answer is fuzzy, the answer is no. Our comparison of LLM brains for Quebec SMBs in 2026 can help clarify.

    What channel do I give the customer for sending images? SMS, WhatsApp, web form, MMS, Messenger: every channel has its own integration logic. The technical complexity hides there, not in the AI itself.

    On budget: expect $30 to $80 per month on top of the standard voice bill, integration and setup included. At Agent IA Vocal, the TECHMA team handles all of that for the client — channel hookup, Law 25 compliance, real-call testing.

    FAQ

    Does the customer have to install an app?
    No. A regular SMS with a photo, or a web link where they drop the image, is enough. The channel feeds into the agent — not the other way around.

    Can the agent see in real time during a video call?
    Not yet in stable production in Quebec. Continuous video streaming is still pricier, slower, and legally riskier. Static photo: yes. Live video: revisit in 2027.

    What happens if the photo is blurry or useless?
    The agent asks for another photo, or escalates to a human. Good systems enforce image-quality rules (resolution, contrast, format). Below threshold, the agent doesn't make up an answer.

    What about privacy?
    The image is processed for the conversation and then purged according to the configured retention policy — often 24 to 72 hours. For sensitive sectors (health, legal, financial) we configure immediate purges or local hosting. The TECHMA team sets that up at install.

    Does multimodal also work in Quebec French?
    Yes. The model doesn't see the language, it sees the image. The voice describing the image speaks your customer's language — Quebec idioms included. Same mechanic as the bilingual code-switching we documented in April.

    So now what?

    April 2026 was a quiet tipping point. For the first time, an affordable AI voice agent can see what the customer is showing it. Not in five years, not in a Stanford lab: today, at $99 a month, on the phone of a Saint-Hyacinthe SMB.

    Does that mean you should rush in? Not necessarily. But it does mean that in 12 to 18 months, the competitors who made the leap will have a measurable edge: fewer dead-end truck rolls, sharper estimates, files opened faster. Ignoring multimodal in 2026 looks a lot like ignoring email in 1998.

    Book 20 minutes with our team to see a live demo on your own use case, or browse the plans starting at $49/month. Multimodal setup is on us — that's what the TECHMA team is for.

    multimodalElevenLabsOpenAIApril 2026Quebec SMBAI voice agent
    Share