Product1 distinct publisher3 min readUpdated
A builder's account decomposes a production voice minute into speech-to-text, tokens, speech synthesis, telephony and media infrastructure, then puts failed calls in the numerator.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
A first-person account published by The Next Web pulls apart the number most voice agent vendors lead with, cost per minute, and argues it reveals little about what a production system actually costs to run [1]. The author, who says they built and operated a voice agent for enterprise use, counts five layers under that single figure, each metered on a different basis [2][3].
Speech-to-text is typically priced per audio minute [4]. LLM inference is priced by tokens, not by time [5]. Text-to-speech is billed by characters, tokens or the volume of speech generated, and becomes a large variable cost when the agent talks a lot [6]. Telephony and transport charge against connected time, varying by provider, call type and location [7]. Real-time voice infrastructure, meaning media servers, orchestration, compute, session rate, logging and monitoring, is folded into per-minute pricing by managed platforms and turns into engineering cost if you build your own [8]. Which is why, as the piece puts it, a headline of $0.05 per minute means little until you know what is inside it [9].
The example that carries the argument is a 35-minute interview-style call in which the participant speaks for 20 minutes and the agent for 12, with pauses, interruptions and turn transitions taking the remainder [10]. That leaves roughly three minutes of neither party talking [1]. Speech-to-text mostly sees the participant's speech, TTS sees the agent's, and telephony and infrastructure may meter all 35 minutes [11]. Agent speech is about a third of that connected time [2].
The model sees another axis: accumulated context [12]. Early on it processes system instructions and a few exchanges; after 20 turns the same length of reply may carry earlier answers, recent dialogue and tool outputs, so a later turn can cost more than an earlier one even when both take the same time to speak [13]. Summarising old dialogue, dropping stale information, retrieving only what is relevant and caching all slow that growth, with the obvious trade: prune too hard and the agent forgets something it needed, keep everything and token usage climbs [14]. Two calls of identical duration diverge here. Five clear answers is cheaper than the same clock spent interrupting, asking for clarification and returning to the original question, because the second call generates more turns [15].
Waste also sits in the gaps. Silence is not charged by speech recognition but is charged by telephony and session infrastructure for as long as the connection stays open [16]. An interruption stops the agent mid-response after TTS has already generated audio nobody hears [17]. Slow responses make people repeat themselves or start a new sentence, which adds turns, processing and connected time [18]. Faster models, better turn detection and more engineering raise direct cost, and the account's position is that they are cheaper than the extra turns a slow system provokes [19].
Then the part that breaks per-minute budgeting outright. If a system needs 108 attempts to produce 100 successful outcomes, the eight failures have already consumed transcription, tokens, generated speech, telephony and infrastructure before ending, and each retry starts another set of costs [20]. That is roughly 7 percent of attempts billed against no result [3]. The working formula offered puts the voice stack, failure and retry overhead, human handling, and evaluation and operations over successful outcomes [21]. It also notes that when in a call the failure happens matters [22].
Two things are worth pressing on. Whether a per-minute quote includes media infrastructure or quietly reassigns it to your engineering budget [8], and whether internal reporting can express cost per successful outcome at all rather than cost per connected minute [21]. Until failed attempts and retries land in the numerator, the number a team optimises is not the number it pays [20][21].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Most voice agent pricing starts with a single number, voice agent cost per minute, but that figure does not reveal much about the real cost of running a production system.
A production voice agent usually has five main cost layers: speech-to-text, LLM inference, text-to-speech, telephony and transport, and real-time voice infrastructure.
Speech-to-text converts incoming audio to text and is typically priced per audio minute.
LLM inference charges by tokens, not by time; longer calls can cost more as the model processes instructions, conversation history and tool results.
Text-to-speech is usually priced by characters, tokens or the amount of speech generated, and can become a high variable cost if the agent talks a lot.
Telephony and transport add charges based on connected time, depending on the provider, call type and location.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source practitioner framework, no measured data
One publisher, one first-person essay. The structural claims about how each layer is metered are stated clearly and are internally consistent, and the two derived figures follow arithmetically from the author's own examples. But every number in the piece is illustrative (a 35-minute call split 20/12, 108 attempts for 100 outcomes), no provider, price sheet, invoice, benchmark or deployment telemetry is cited, and the first-hand enterprise experience underpinning the argument is self-reported and unverifiable in the supplied text. That supports a reasoning framework, not a measured cost finding.
No adoption signal in supplied material
The supplied source reports no release, deployment, benchmark, pricing change, license change or usage disclosure. It names no vendor, platform, customer or call volume, and offers no evidence that any team has moved to per-outcome pricing. There is nothing to measure adoption against, and inferring uptake from a framework essay would be guessing.
Mildly overstated generality, deflationary in intent
The piece argues in the deflationary direction - it tells readers a low headline per-minute price is misleading and sells no product - which keeps the gap small. The residual overstatement is in reach rather than pitch: 'usually five main cost layers' and the crisp illustrative figures (a 35-minute call, 108 attempts for 100 outcomes, a $0.05 minute) read as industry regularities and near-empirical rates while resting on one unverified first-hand account with no vendor pricing or deployment data behind them. Slightly positive, not substantially inflated.
Author interest undisclosed
The supplied text gives no author identity, employer or commercial-interest disclosure, and names no vendor, platform or product that the argument could favour or disfavour. A first-person builder essay on a commercial tech publication may or may not carry an underlying commercial stake in voice infrastructure; the material supplied does not establish either way, so no incentive score is asserted.
Coherent but uncorroborated and unmeasured
Confidence is limited by cluster shape: a single publisher, a single item, no adoption evidence and no incentive disclosure. What raises it above the floor is that the claims are mostly definitional or mechanistic, are uncontested within the material, and hang together logically, with the derived figures checkable against the source's own numbers. The right posture is to treat the framework as a hypothesis to instrument, and to hold the quantitative examples loosely.
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
product
Cinemas, classrooms and ICE: smart glasses now need a venue-policy contingency1 distinct publisher
build
The demo-best voice engine finished last: 12,247 calls argue for buying on completion rate1 distinct publisher
product
Anthropic wants a bigger raise than SpaceX and will not say what it is worth2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026