Skip to content

Product1 publisher3 min readPublished

Five meters, one minute: why voice agent budgets should be priced per outcome

A builder's account decomposes a production voice minute into speech-to-text, tokens, speech synthesis, telephony and media infrastructure, then puts failed calls in the numerator.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Five meters, one minute: why voice agent budgets should be priced per outcome
Photo: thenextweb.com

What happened

  • Most voice agent pricing starts with a single number, voice agent cost per minute, but that figure does not reveal much about the real cost of running a production system.
  • The author says they discovered this firsthand while building and running a voice agent for enterprise use.
  • A production voice agent usually has five main cost layers: speech-to-text, LLM inference, text-to-speech, telephony and transport, and real-time voice infrastructure.
  • Speech-to-text converts incoming audio to text and is typically priced per audio minute.
  • LLM inference charges by tokens, not by time; longer calls can cost more as the model processes instructions, conversation history and tool results.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

A first-person account published by The Next Web pulls apart the number most voice agent vendors lead with, cost per minute, and argues it reveals little about what a production system actually costs to run [1]. The author, who says they built and operated a voice agent for enterprise use, counts five layers under that single figure, each metered on a different basis [2][3].

Speech-to-text is typically priced per audio minute [4]. LLM inference is priced by tokens, not by time [5]. Text-to-speech is billed by characters, tokens or the volume of speech generated, and becomes a large variable cost when the agent talks a lot [6]. Telephony and transport charge against connected time, varying by provider, call type and location [7]. Real-time voice infrastructure, meaning media servers, orchestration, compute, session rate, logging and monitoring, is folded into per-minute pricing by managed platforms and turns into engineering cost if you build your own [8]. Which is why, as the piece puts it, a headline of $0.05 per minute means little until you know what is inside it [9].

The example that carries the argument is a 35-minute interview-style call in which the participant speaks for 20 minutes and the agent for 12, with pauses, interruptions and turn transitions taking the remainder [10]. That leaves roughly three minutes of neither party talking [1]. Speech-to-text mostly sees the participant's speech, TTS sees the agent's, and telephony and infrastructure may meter all 35 minutes [11]. Agent speech is about a third of that connected time [2].

The model sees another axis: accumulated context [12]. Early on it processes system instructions and a few exchanges; after 20 turns the same length of reply may carry earlier answers, recent dialogue and tool outputs, so a later turn can cost more than an earlier one even when both take the same time to speak [13]. Summarising old dialogue, dropping stale information, retrieving only what is relevant and caching all slow that growth, with the obvious trade: prune too hard and the agent forgets something it needed, keep everything and token usage climbs [14]. Two calls of identical duration diverge here. Five clear answers is cheaper than the same clock spent interrupting, asking for clarification and returning to the original question, because the second call generates more turns [15].

Waste also sits in the gaps. Silence is not charged by speech recognition but is charged by telephony and session infrastructure for as long as the connection stays open [16]. An interruption stops the agent mid-response after TTS has already generated audio nobody hears [17]. Slow responses make people repeat themselves or start a new sentence, which adds turns, processing and connected time [18]. Faster models, better turn detection and more engineering raise direct cost, and the account's position is that they are cheaper than the extra turns a slow system provokes [19].

Then the part that breaks per-minute budgeting outright. If a system needs 108 attempts to produce 100 successful outcomes, the eight failures have already consumed transcription, tokens, generated speech, telephony and infrastructure before ending, and each retry starts another set of costs [20]. That is roughly 7 percent of attempts billed against no result [3]. The working formula offered puts the voice stack, failure and retry overhead, human handling, and evaluation and operations over successful outcomes [21]. It also notes that when in a call the failure happens matters [22].

Two things are worth pressing on. Whether a per-minute quote includes media infrastructure or quietly reassigns it to your engineering budget [8], and whether internal reporting can express cost per successful outcome at all rather than cost per connected minute [21]. Until failed attempts and retries land in the numerator, the number a team optimises is not the number it pays [20][21].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories