The real margin on reselling AI voice comes from the spread between a wholesale cost stack that runs in cents per minute (carrier plus LLM plus text-to-speech) and a resale price billed as a flat monthly subscription per seat or per line, and that spread stays wide as long as the agency, not the platform, controls the client contract and the overage terms. There isn't one industry-standard margin percentage, because the two sides of that spread move independently: usage costs are variable and small, resale pricing is usually fixed and set by the agency. Get the structure wrong and a client who calls in heavily can quietly erase a month of profit on that account alone.
What's actually in the cost stack?
Three line items make up the wholesale side: carrier minutes for the actual phone call, LLM tokens for the conversation reasoning, and text-to-speech characters for the voice output. Carrier pricing is public: Twilio's voice pricing page lists per-minute rates that vary by destination and route, typically a fraction of a cent to a few cents per minute for standard PSTN traffic. LLM cost depends on which model handles the call, and providers publish this openly too, OpenAI's pricing page shows per-token rates that differ by an order of magnitude between a lightweight model and a frontier one. That gap matters for margin: a platform that lets you route calls to a cheaper model for routine bookings and a stronger one only for complex conversations keeps the wholesale side down without touching call quality. Being able to choose across multiple LLM and TTS providers, rather than being locked to one vendor's pricing, is one of the few places an agency can actually move the cost line instead of just accepting it.
How does the resale price get set, and why does that decide the margin?
Most agencies bill clients a flat monthly rate per seat, per phone line, or per bundled minute allotment, not a live pass-through of usage. That's the mechanism that creates margin: usage is metered in fractions of a cent, resale is billed in whole dollars per month. The spread holds as long as the client's actual usage stays under the volume the flat price was built to absorb. /agency-pricing lays out the tiers agencies build on top of, and the companion piece on pricing white-label voice AI without eating markup walks through where agencies commonly set that ceiling too low and give the margin back on the accounts that actually use the product the most.
Why isn't there a single margin number that applies to every agency?
Because the variables aren't fixed. A dental-vertical client answering appointment calls behaves very differently, usage-wise, from an outbound sales dialer running hundreds of calls a day. An agency using BYOK to plug in its own LLM and carrier credentials carries the wholesale cost directly and can shop rates, which changes the math versus a platform that bundles usage into the subscription. Client count matters too: the first few accounts absorb fixed platform and setup costs before any of it is margin, so the real number is closer to a curve that improves with volume than a static percentage anyone can quote up front.
Where do margins actually get eaten in practice?
The two most common leaks aren't cost overruns, they're structural. First, flat-rate plans without a usage ceiling: a client that runs far more call volume than the plan assumed turns a profitable account into a break-even one, which is why usage caps and overage terms belong in the client contract, not left to trust. Second, losing control of the underlying account. Some voice AI platforms tie carrier and billing setup to the vendor rather than the reseller, which means the agency is exposed if pricing or terms change underneath them, worth checking directly against a given vendor's own docs before building an economics model on top of it, as with any Vapi comparison an agency runs before choosing a platform. Per-tenant carrier isolation, where each client's calling infrastructure is kept separate rather than pooled, is one of the mechanisms that keeps that risk from compounding across a growing client roster.
Does the margin change once an agency scales past a handful of clients?
Generally yes, and in the agency's favor. Onboarding and setup work per client shrinks as the process gets repeatable, and usage costs stay near-flat per minute regardless of client count, so the fixed costs that eat into the first few accounts' margin get spread thinner as the roster grows. This is the core argument for treating white-label voice as a subscription line of business rather than a one-off project fee, covered in more detail on /agencies. It's also why the actual dollar margin an agency nets in year one looks different from year two even if the resale price per client never moves.
What should an agency actually model before committing to a resale price?
Three inputs: expected average call volume per client (not the ceiling, the realistic average), the wholesale cost per minute across whichever carrier, LLM, and TTS combination is in use, and the overage terms that kick in when a client exceeds what the flat price assumes. Skipping the third one is the most common mistake, because it's the only input that protects the margin on the outlier accounts rather than just the average ones. An agency that models all three before setting a resale tier on /agency-pricing is pricing off real math instead of guessing at a round number and hoping usage stays polite.