Voice AI Pricing in India Keeps Falling. What If It Doesn’t Matter Who Wins?

The lowest published voice AI rate in India has changed hands several times this year, two rupees a minute, then one and a half, with no clear sign of where it settles. Published rates across the market span roughly ₹1.50 to ₹20 a minute, and the vendor sitting at the low end hasn’t stayed there long enough to call it a fixed market price. This is a subsidy war still being fought, not a market that has found its floor.
The read making the rounds right now is broadly correct: this isn’t a winner-take-all market the way early e-commerce or ride-hailing were. Voice AI orchestration is largely software wrapped around commodity infrastructure, speech-to-text, inference, speech synthesis, most of it rented rather than owned. Bolna, one of the venture-backed entrants building specifically in this layer, describes its own product as an orchestration layer connecting different voice models for enterprises, an honest description of what this category actually is. Commoditized infrastructure like that does not produce durable pricing power. When the venture capital funding a ₹1.50 rate runs out, the vendor either shuts down or reprices five times over to survive, and either way, a business case built on today’s voice AI pricing stops holding.
There’s a comparison attached to this argument, that eventually someone “pulls a Jio” and forces a shakeout the way Reliance Jio did in Indian telecom. That comparison deserves a closer look rather than a quick dismissal, because Reliance has already made close to this exact move once. It acquired a majority stake in Haptik in 2019 specifically to get into conversational and voice AI, and at its 2026 AGM it went further, unveiling Jio TeleFrame and a network-embedded voice agent built on a telecom subscriber base of 450 million-plus, an asset no per-minute software vendor in this category owns or can rent. That’s a real structural advantage, not a marketing claim.
Even so, seven years after that acquisition, the standalone per-minute price war among software vendors has kept fragmenting anyway, a fresh “lowest rate in India” claim still surfaces from a new entrant on a regular cycle, largely untouched by Haptik’s presence. Owning the network hasn’t been enough on its own to settle this category the way it settled mobile data. So there are genuinely two ways this plays out: consolidation around a distribution-advantaged incumbent, or continued churn among venture-backed software vendors. The exposure for an enterprise is identical either way. Wired to a software vendor that gets undercut or shuts down, an enterprise is stranded. Wired to a vendor that gets acquired, deprioritized, or folded into someone else’s roadmap, it’s stranded just the same.
Recalculating the ROI with realistic margins is the right instinct against that risk. It is not the whole fix, because it only addresses one exposure and there are at least three.
The first is the subsidy risk already described. The second sits inside the voice AI pricing model itself, and it doesn’t require a single vendor to ever reprice. Most of this market still bills per minute, and per-minute billing pays for duration, not resolution. One audited deployment showed what happens when a script gets tightened rather than a price renegotiated: average handle time dropped from 3.1 minutes to 2.2, completion rates held steady, and the monthly bill came down by close to 29 percent, savings a per-minute biller has no built-in reason to go find. That tension is structural. It survives even in a world where voice AI pricing never moves again.
The third exposure has nothing to do with price at all. A vendor can become non-viable overnight because of where customer voice data is allowed to live, not because of what it costs. Sarvam, one of the more full-stack players in this category, is already building around that risk explicitly, offering private cloud, on-premise, and fully air-gapped deployment with SOC 2, ISO 27001, and DPDP compliance built in for regulated Indian sectors. That’s a tell. If compliance-driven stranding is a real enough concern that a vendor is architecting an entire deployment model around it, it is a real enough concern for every enterprise in this category to plan for, not just the ones already worried about it.
None of that is a knock on the vendors doing the modeling and the voice work well. ElevenLabs, by independent review, is arguably the strongest conversational engine in the category, sub-300-millisecond latency and natural-sounding speech across dozens of languages. The same review is just as clear about the boundary: not built as a call center platform, no native ticketing, no CRM sync, no agent supervision tooling, out of the box. That’s not a flaw. It’s a scope. The gap is exactly the layer this piece has been describing, and closing it was never the job these vendors set out to do.
All three risks share a root cause. None of them are about whether a particular model, engine, or API is good. They’re about what happens to the operation underneath it when the vendor, the price, or the regulation changes, and in a category this early, something reliably will. The fix isn’t a more conservative forecast. It’s an architecture that doesn’t need one, where the model is treated as the one part of the stack that is supposed to change, and everything else, oversight, quality, memory, the compliance record, stays constant regardless of what gets swapped underneath.

On Vitos, that separation is a configuration screen, not a migration project. LLM, voice provider, transcriber, and knowledge base each sit as independently swappable layers, with a live connection status across the whole stack.
Every layer of the stack is swapped on its own, without touching the others.
A change here doesn’t go straight to every caller. A new choice runs first in sandbox, or against a defined control group, measured against the same resolution and quality benchmarks the outgoing choice was held to. Only a validated change gets promoted live. That’s what makes “not being stranded” an operational fact rather than a line in a pitch deck.
Which raises the obvious next question. If the model isn’t where the durable value sits, where is it? By our latest count, Kapture runs across an estimated 1,000-plus enterprise customers, with nearly 200 AgentOS deployments live this year and an estimated 2 billion-plus agentic AI interactions processed since last year. But the number that matters more than any of those is ten, as in ten years spent inside the actual workflows of large Indian retail and BFSI brands, watching how their interactions with customers work, where those interactions break, and what a resolution really looks like in each of those businesses. That experience turned into templates and workflows an enterprise, or a smaller team without one, can pick up and customize instead of building an agent from a blank canvas, with forward-deployed engineers on hand for the parts that still need building by hand. That’s not a feature. It’s the part of the stack a model swap was never going to touch.
Cheaper voice models are, on their own terms, good news, lower marginal cost and more room to experiment. The only businesses for whom this price war turns into a trap are the ones whose roadmap depends on knowing who wins it. For everyone else, it’s noise, not signal.
Voice AI pricing will keep moving. The infrastructure underneath a CX operation should not have to move with it.
Kapture works with BFSI and retail enterprises across India to build that separation from day one. Get in touch to see it on your stack.
Your Plan. Your Value. Your Growth.
Your business is different – and the pricing should reflect that.
Let’s build a plan that matches your goals, maximizes ROI, and scales with your success.






