You price white-label voice AI without eating the markup by knowing your true per-minute cost (telephony, LLM tokens, TTS, and platform fee) before you quote a client, then building a rate card with a fixed buffer above that cost so a busy month doesn't turn a profitable client into a break-even one. The mistake most agencies make isn't picking the wrong number, it's quoting a flat monthly price without modeling what happens when usage doubles.
What actually makes up your cost stack?
Every voice AI call has at least four cost components stacked on top of each other: the telephony carrier leg (inbound/outbound minutes), the LLM inference (tokens in and out per turn), the text-to-speech synthesis, and whatever your platform charges you as the reseller. Telephony alone varies by carrier and route, Twilio's public voice pricing page shows per-minute rates that shift by country and whether the call is inbound or outbound, so a "typical call cost" only means anything once you've picked your carrier and traffic pattern. LLM cost is usage-based too: OpenAI's API pricing page lists per-token rates that differ by model, and a longer, more conversational call burns more tokens than a quick "what are your hours" exchange. If you don't know these numbers cold for your own stack, you're guessing at your margin, not calculating it.
This is also where BYOK (bring your own key) changes the math. Running your own LLM and TTS keys instead of paying platform markup on inference can meaningfully change your cost per minute, but only if you're checking that math against your actual call volume rather than assuming savings that don't materialize at low volume. Our BYOK savings calculator is built for exactly that comparison before you commit a client to a rate.
Should you bill per-minute, per-call, or flat monthly?
Each model shifts risk differently:
- Per-minute billing passes usage risk to the client but is harder to sell because SMBs want predictability, not a bill that swings with call volume.
- Per-call billing is easier to explain but punishes you on long calls (a 12-minute troubleshooting call costs the same to you as three 4-minute calls, but you're billing it as one unit).
- Flat monthly with a usage cap is what most white-label agencies land on: a fixed price for up to X minutes or calls, with overage billed at a stated per-minute rate above that. This gives the client budget certainty and gives you a hard ceiling on how much unpaid usage you'll absorb.
The cap is the part agencies skip and then regret. Without one, a single client whose call volume triples in a busy month eats your margin for that entire billing cycle, and you find out after the invoice, not before.
How much buffer do you actually need above cost?
There's no universal number here, and any post that gives you one without showing its math is guessing. What you can do is build your own buffer from your own cost stack: take your worst-case per-minute cost (peak LLM token usage plus your carrier rate plus platform fee), multiply by your expected monthly minutes per client, and that's your floor. Your rate card price needs to clear that floor with room left over for support time, onboarding, and the inevitable client who calls you every week asking to change the greeting script. Agencies that skip this step and price off a competitor's public number end up matching a price built on a cost stack they don't have, which is why comparing your setup against named platforms like Vapi or Retell on infrastructure and support model, not just sticker price, matters before you set your own card. Our agency pricing page breaks out the tiers we built this way so you can see the structure rather than reverse-engineer it.
What's the real edge case that eats margin?
The client who insists on unlimited minutes at a flat rate "because that's what my last vendor did." Named competitors publish their own plan structures, and it's worth reading how they cap or don't cap usage on their own pricing page before you agree to match a client's expectation you can't actually support at cost. If a prospect wants unlimited at a fixed price, that's a signal to either quote a higher flat tier with a real minute ceiling, or walk. Underpricing one client to win the deal doesn't just cost you margin on that account, it sets an expectation the next five prospects will point back to.
The other edge case is outbound. Outbound minutes carry different carrier costs and different compliance overhead than inbound, and lumping them into the same flat rate as inbound reception minutes usually means one side is subsidizing the other. If you're running outbound campaigns for clients, price that usage separately with its own cap, not folded into the same bucket as inbound call handling.
How do you know if your pricing model can actually scale?
Run the math forward, not just for one client but for what happens at 20 or 50 clients on the same tier. A rate card that clears margin comfortably for one SMB client at moderate volume can compress fast if you scale that same tier across a book of clients without re-checking your infrastructure cost at volume. This is also where it's worth being honest with yourself about whether your agency is actually ready to operate this at scale, not just sell it. Our agency readiness scorecard walks through the operational side of that question: support load, onboarding time, and churn risk, all of which affect your real margin more than the number on your rate card does.
What's the follow-up question worth asking before you set your card?
Ask yourself what happens the month a client's call volume triples and whether your current pricing structure has an answer for that scenario already built in, or whether you'd be renegotiating mid-contract. If you don't know the answer, that's the gap to close before you quote your next client, not after. Start from your actual cost stack, build the buffer and the cap in from day one, and check your full pricing structure against the tiers on our pricing page before you finalize your own.