How to Measure Your AI Voice Agent's Real ROI in 14 Days (The Deflection Test Vendors Don't Want to Show You) | Agent IA Vocal
    Back to blog
    How-to Guide8 min readApril 21, 2026

    How to Measure Your AI Voice Agent's Real ROI in 14 Days (The Deflection Test Vendors Don't Want to Show You)

    The 14-day test that reveals your AI voice agent's true deflection rate. A/B protocol, ROI formulas, honest benchmarks for Quebec SMBs — no vendor demo bluff.

    MA

    Masdouk Adelakoun

    Cofondateur & CTO

    How to Measure Your AI Voice Agent's Real ROI in 14 Days (The Deflection Test Vendors Don't Want to Show You)

    Your vendor promised an 80% resolution rate in the demo. Three weeks into the rollout, your AI voice agent is escalating 7 calls out of 10 to your receptionist — who is now busier than before.

    This isn't a bug. It's that you measured the wrong thing. Or more precisely: the vendor showed you the metric that flatters the sales deck, not the one that protects your ROI.

    Since OpenAI's GPT-realtime went generally available in late 2025 and ElevenLabs rolled out multi-agent workflow evaluations, voice agent demos have started quoting even more impressive numbers. But in production, inside a real Quebec SMB, with bilingual calls, regional accents, and callers who hang up at second three — the real number collapses. Sometimes by half.

    The good news: there's a 14-day test, simple enough for any owner-operator to run, that tells you whether your agent actually delivers. And the metric that decides is the deflection rate. Not the CSAT. Not the success rate. Not the NPS. The deflection rate.

    Why 9 SMBs out of 10 miscalculate their voice AI ROI

    The classic trap: the vendor flashes a number — "82% of conversations resolved without a human" — pulled from a filtered sample. Calls hung up in under 5 seconds? Excluded. Calls transferred to voicemail? Counted as "resolved." Calls where the caller gave up because the agent couldn't parse a Saguenay accent? Invisible in the dashboard.

    An owner we talked to last week thought he was hitting 73% deflection. After audit, his real number was 34%. The gap? The vendor counted every call that didn't create a ticket as a "win" — including calls where the client called back 20 minutes later from a different number.

    Real deflection rate, by contrast, does not lie. It's the percentage of calls your agent handled from start to finish — without transfer, without a follow-up callback, without a delayed escalation, without measurable dissatisfaction in the next 72 hours.

    What a "deflection rate" really means — the audit-proof definition

    The formula that matters:

    Real deflection rate = (Calls fully handled by the AI AND not followed by a human callback within 72h) ÷ (Total inbound calls)

    Three anchor words:

    • Fully — not "80% of the conversation." The caller leaves satisfied on the AI line.
    • Not followed — if the same caller calls back within 3 days about the same issue, that's not a deflection.
    • 72 hours — the window that eliminates noise callbacks without diluting the measure.

    No vendor volunteers this definition. It drops the reported number by 15 to 25 points. But it's the one that predicts whether your ROI is real or theoretical.

    The 3 metrics to track in parallel (don't fixate on deflection alone)

    Deflection rate taken in isolation can deceive. An agent that hangs up on callers hits 100% deflection. Here are the three you read together:

    1. Net deflection rate (as defined above). Honest target: 35-55% in week one, 55-70% after 90 days of tuning. Above 75% without a human backup, that's a red flag — audit your measurement.

    2. Containment rate. Percentage of calls never escalated, even mid-conversation. Different from deflection: a call can be "contained" (never transferred) but not "deflected" (caller unhappy, calls back next day). Target: containment should exceed deflection by no more than 5-10 points. Larger gaps mean your agent escalates poorly.

    3. First call resolution (FCR). The real judge. Caller calls, AI resolves, caller never calls back. Quebec SMB target: 50-65% in steady state. Below 40% points to a qualification or knowledge-base problem.

    These three numbers read together give you a picture neither vendor nor demo can doctor.

    The 14-day deflection test: exact protocol

    This is the protocol we run for every client. You can run it yourself. It requires no fancy dashboard — just a spreadsheet and a bit of discipline.

    Days 1-3: establish the human baseline

    Before measuring anything on the AI, measure what you have today. Record over 3 days: inbound call volume, pickup rate by your team, average handling time, abandoned callbacks. Skip this step and you have no reference point — so no calculable ROI.

    A typical Quebec SMB with 15-50 employees sees 40-120 inbound calls per day, 55-70% pickup rate, and 8-15% callbacks lost within the same day. Those are your baseline.

    Days 4-10: A/B deployment with 50/50 routing

    Do not flip 100% of calls to the AI on day one. Route 50% to the agent and 50% to your human team — ideally by time slot or source number. Compare both cohorts on the same questions:

    • How many calls resolved without further contact within 72h?
    • How many triggered a callback (same caller, within 3 days)?
    • How many hung up in under 10 seconds?

    Watch for one frequent bias: if you route "easy" calls to the AI and "complex" ones to the team, you artificially inflate AI performance. Routing must be random or keyed on a neutral field (e.g., even/odd phone number).

    Days 11-14: final measurement and ROI calculation

    Compile the numbers from days 4-10. Apply:

    Monthly ROI = (Real deflection × Monthly volume × Cost per human call) − (Monthly AI agent cost + supervision cost)

    A concrete example for a 4-practitioner Quebec dental clinic, 80 calls/day:

    • Monthly volume: 1,760 calls (22 working days)
    • Measured deflection: 48%
    • Cost per human call (lost time, callbacks included): $4.20
    • Gross savings: 1,760 × 0.48 × $4.20 = $3,548/month
    • AI agent + supervision cost: $650/month
    • Net ROI: $2,898/month, or roughly $34,780/year

    If your test shows a negative or marginal ROI, two options: renegotiate tuning with the vendor, or switch tools. What you must not do is keep the agent out of inertia, telling yourself "it'll improve." Agents that miss their numbers after a clean 14-day A/B test never recover — they get worse, because the human team disengages in parallel.

    The Quebec-specific trap: FR↔EN code-switching

    One detail that wrecks generic-agent stats in Quebec: callers mix French and English in the same sentence. "J'aimerais book un rendez-vous pour ma fille, she's been having mal à la gorge depuis vendredi." Agents tuned on standardized French fall over. OpenAI's new GPT-realtime model handles this mid-sentence switching — but only if the vendor's configuration explicitly enables it.

    In your 14-day measurement, separate the deflection rate on monolingual vs. bilingual calls. The gap reveals the real quality of the chosen model. A gap wider than 20 points between the two cohorts means reconfiguration is needed — see our piece on FR-EN code-switching traps in production.

    What to demand from your vendor before signing

    Before running this test, ask your vendor in writing for three things. If they refuse any one of them, you have your answer.

    1. Access to raw 72-hour post-call logs — to verify callbacks and measure real deflection, not dashboard deflection.
    2. Ability to run 50/50 routing during the A/B phase — some vendors only allow 100/0 or 90/10, which makes clean measurement impossible.
    3. An exit clause if real deflection stays below 35% at day 14 — no cancellation fees. A confident vendor accepts.

    For the full checklist, see our reference on 5 scenarios to run before signing with a voice agent vendor. And if you want to know where fees explode once the contract is signed, we already dissected the 7 hidden fees that inflate vendor bills.

    How TECHMA runs the tracking for you

    Measuring real deflection requires wiring up call logs, a 72-hour callback tracer, a clean A/B split — and interpreting the numbers without being bluffed by flattering metrics. Our team handles it end-to-end: A/B routing setup, 72-hour tracking instrumentation, weekly deliverable with the three key metrics, and tuning recommendations if numbers slip.

    The 14-day test becomes your safety net, not another project you have to carry. And most importantly: two weeks in, you know whether the agent belongs in your SMB — backed by numbers no vendor can contest.

    The real question isn't "does it work"

    The real question is: "Do I have the right metric to know?" Vendors know that a flattering CSAT and a vague "resolution rate" are enough to sell renewals. Net 72-hour deflection, measured on a clean A/B split, over 14 days, takes that away. Either the agent performs, or you see it — and you decide.

    For SMBs still hesitating between an inbound and an outbound agent before they even measure, our inbound vs. outbound decision framework clarifies the call up front — because you don't measure the same rate for the same use case.

    Cobbai's industry benchmarks place average deflection between 20 and 40% across sectors — nothing magical. But with the right tools and the right measurement discipline, well-supported Quebec SMBs routinely clear 55% after 90 days. The difference isn't the vendor. It's the rigor of the initial test.

    One last reality check: the tooling is finally mature enough

    Two years ago, running this test on a real SMB wasn't worth it — the underlying models simply couldn't hold a natural conversation past the first branching point. That barrier is gone. Between ElevenLabs' multi-agent workflow and evaluation scoping rolled out in Q1 2026 and OpenAI's GPT-realtime SIP connector, the ceiling on what a well-configured Quebec agent can handle in production has moved meaningfully. Which is exactly why a 14-day test today tells you something it couldn't tell you in 2024: whether your configuration, on your call mix, clears the bar. Not whether the technology in the abstract works — that's already settled.

    Run the test. Get the numbers. Decide with data, not with the demo deck.

    Share