The Top 4 Buying Models for Voice Agents in CX

Companies have spent well over $100 billion a year on customer experience with no detectable return, according to the American Customer Satisfaction Index. The national ACSI score sat at 76.7 in Q1 2026, the same level it held in 2013, while customer complaints surged 16% to record highs. By Q2 2026, the index had declined sharply again.

Decades of investment, and yet customer anger is still growing. Every major CX technology wave was sold on cost per contact and delivered exactly that. IVR routed calls so fewer humans had to answer them, and customers learned to mash zero. Scripted chatbots gave scripted answers to unscripted problems. Self-service portals moved the labor onto the customer and counted it as a win. Each one worked as designed. The design was to spend less, not to resolve more.

Voice AI is now arriving into that same procurement reflex, and the build-versus-buy debate is where the reflex shows up first.

The short answer: build vs. buy is the wrong axis. The market has sorted into four buying models, and they differ less on price than on opportunity cost: what each one quietly takes off the table later. Day-one capability no longer separates them, because most credible vendors now clear the same bar. What separates them is where you can still get to in month 12.

Day one shouldn’t be the deciding factor

Run a bake-off today, and a lot of vendors will clear the same bar, most of them well. Replace the IVR. Connect to your CRM and ticketing. Transfer to a human with context attached. Hold concurrency at peak. Report containment and AHT.

Two years ago that list was a real differentiator. It’s close to table stakes now, which is good news for anyone buying, but it also means a strong demo tells you less than it used to.

That matters more in voice than in most software purchases, because of what you’d have to move if you chose wrong. By month 12 you’ve built conversation logic, escalation rules, tool calls into your systems, and evals that encode what good sounds like for your customers. That work is the accumulated understanding of your own operation. Rebuilding it on a different provider isn’t a migration. It’s doing the year again.

The four buying models

1. Build it yourself

Suits: Companies with voice as a core product surface and engineering capacity to match.

You own everything, and your engineering cycles go to concurrency, failover, telephony carriers, and integrating every new model as it ships. The agent itself is a fraction of that work, and teams discover the ratio late, usually when a demo that ran one call at a time turns into a quarter of unscoped infrastructure work.

The cost that gets missed is maintenance rather than construction. Carriers change, transcription providers change, models are deprecated on someone else’s schedule. That’s a standing team, not a project. Build only if the voice stack is a differentiator in itself and you’re prepared to fund it in year three as heavily as year one.

2. Verticalized point solutions

Suits: Teams in a well-templated industry whose processes genuinely resemble the template.

Incumbent contact center suites and vertical AI vendors ship an agent pre-shaped for your sector. Fast, because the decisions have already been made for you, which is also the problem. You have no real power to shape it, so wherever your business logic diverges from the template, the template wins and you adapt your operation to the software.

That trade is fine when your process is standard and the goal is deflection. It gets expensive when the thing that differentiates you commercially is precisely the thing the template flattens. The exit cost is high because almost nothing you configured is portable.

3. Horizontal agent apps

Suits: CX teams that want a working agent quickly and expect to stay in the middle of the ladder.

CX-specific agent applications, more configurable than a vertical point solution and genuinely quick to something live. You shape the agent in their format, which means configuration happens inside their abstractions and those abstractions are your ceiling.

Model choice, prompt-level control, and eval design usually sit on their side of the line, so improvements arrive on their roadmap rather than yours. Two consequences follow. Your cost base tracks their pricing rather than the underlying model market, which has been falling fast. And your call data trains their system, which means it also funds what they build for your competitor.

4. Open platform

Suits: Teams that intend to move up the ladder and want the iteration loop in-house.

You own the prompts, models, logic, evals, and data, while the vendor runs telephony, orchestration, concurrency, and compliance underneath. You give up building infrastructure and keep every layer that shapes the outcome.

The cost is real: this expects someone on your side to own the agent, and it isn’t the right answer for a team that wants a finished product on day one. What you get for that is that changes you ship yourself compound into IP built from your own call data, while on the other approaches, the ones you have to request, come back as delays on someone else’s roadmap.

Too little control caps what you can achieve. Too much, and the cost of ownership eats the return before you reach the rungs worth reaching. The real question isn’t build or buy. It’s which layers you need to own to move the metric you care about, and whether the vendor will let you own them.

Which model you need depends on your goals

Grade a voice agent by what the interaction is worth, and the answer to the question above falls out of it.

Deflect the FAQ. Answer, route, hand off. Measured by deflection, containment, and cost per contact. The question is how cheaply you handled it.

Understand, then hand off. The agent knows what needs to happen but can’t do it. Measured by transfer rate and whether context survived.

Resolve it end to end. Diagnose, act inside your systems, confirm with the customer. Measured by first-contact resolution.

Save the cancellation. Handle the objection, keep the account. Measured by retention and lifetime value.

Sell and expand. Book, upsell, follow up until it’s done. Measured by conversion and revenue per call.

Return per interaction climbs as you go up, and so does the share of your own business logic the agent has to execute. The bottom two rungs are where your entire shortlist competes, because those rungs demand the least access to your systems and the least configurability from the vendor. All four models can deliver them.

Above that, the field narrows fast. A containment-optimized agent and a retention-optimized agent aren’t the same product at different maturity levels. Containment needs a good demo. Saving a cancellation needs you to be able to rewrite objection handling on a Tuesday because of something a customer said on a Monday.

What it looks like when the loop is owned

Three deployments worth studying, each of which ended up somewhere its business case didn’t describe.

A major home security brand moved 100% of inbound support traffic to voice agents inside two weeks. Time on call fell 50% and cost per minute fell 70%, while CSAT kept improving. The detail that matters is who does the tuning: their CX team iterates on the agent directly, without engineering tickets, so seasonal surges no longer mean seasonal hiring.

The largest car marketplace in Latin America started in inbound support and now runs financing, trade-ins, and delivery end to end. Revenue in its Mexico market is up 200% while serving twice the customers, and NPS is up 20 points. Business teams there build new evaluations in about five minutes. Their calls got longer, not shorter, which was the right direction for the job.

A health insurance brokerage running a 400-plus agent operation put four engineers on it and matched the call volume of a 50-person call center in one week, with transfer-to-close up 25%. Sales training leaders, not engineers, now modify the conversation flows.

In each case, the people closest to the customer could granularly change the agent without going through anyone else.

The metric you sign for isn’t necessarily the metric you’ll manage tomorrow

Most deployments start with a narrow, defensible target written into the business case. Deflect the top 20 call drivers. Handle overflow so we stop hiring for seasonal peaks. That’s usually the right place to start.

Then it works, and the goal moves. Six to twelve months in, someone watches the agent resolve a billing issue end to end and asks why it can’t handle the cancellation that follows. The number the deployment was justified on stops being the number anyone cares about, well after the contract is signed.

For the car marketplace above, time on call eventually stopped being a cost line and turned into a sales input, a reversal that would read as failure against the original success criteria. The home security brand went the other way and added CSAT alongside containment, tuning for a quality metric using a system bought on efficiency.

A vendor selected against a containment target will do well against it for as long as that’s the target. When the target moves, you find out whether what you bought can move with it, or whether moving means a migration, a renegotiation, or a roadmap request sitting behind another customer’s.

Day-one demos are getting harder and harder to differentiate. Evaluate on the rate of change instead: not what the agent does now, but how fast, how independently, and how granularly your own team can change it. Pick the vendor whose ceiling you can’t hit inside the contract term, because a good-enough agent you control can be improved, while a great agent you don’t control quietly stops fitting the job as soon as the target changes.


Guest post written by Ryan Ratner, Product Marketing at Vapi

Vapi is an open voice AI platform that gives enterprise teams full control over the models, prompts, logic, and data behind their voice agents.

To learn more about evaluating voice AI for customer experience, visit vapi.ai.