Product3 publishers3 min readPublished
Modulate raises $25 million to sell voice-agent builders a separate audio-analysis layer
Modulate raised $25 million to put models that analyze raw call audio, not transcripts, in front of more developers. Its pitch to teams running voice agents is a separate layer that spots cloned voices and grades how agents handle callers.
The Product Desk · Product desk

What happened
- Future Ventures led the round, with returning investors Hyperplane and Lakestar also taking part.
- Velma, Modulate's main platform, picks from more than 100 small specialized audio models for each task and blends their results.
- Health care institutions use the models to screen out callers who impersonate staff with deepfake voices.
- The money goes to new SDKs, APIs and industry-specific models, plus hiring in research, engineering and developer relations.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision A team with a voice agent in production has to decide whether call audio should also flow to a second vendor, with its own contract and privacy review, for checks its voice stack does not run.
- cost The only public price covers batch transcription, so the cost of the deepfake and agent-grading layer Modulate is actually pitching cannot yet be planned from published numbers.
- exposure Hospitals and call centers that put deepfake screening on live calls depend on a vendor with 40 to 45 employees, so its support capacity belongs in their risk review.
- capability If the promised SDKs and APIs ship, teams building voice products can buy audio analysis off the shelf instead of training their own models.
Carter Huffman's example of a failed call is a polite one. "I think when companies think of emotion analysis, they think if the customer was neutral or positive, the call was a success, and if the customer was negative, the call was a failure. But actually, many times people will be polite even to, like, AI agents or bots. Right. And they won't come across as angry, but they'll be very dissatisfied," the Modulate chief executive told TechCrunch [15][14]. By Huffman's account, teams tell themselves that a caller who stays neutral or positive had a good call. What users actually do is stay courteous to the bot and leave unhappy. Modulate's models work on the raw audio and read tone and intent as signals. They can flag a caller losing patience with a voice agent while the call is still going [3]. The customers who check their agents this way are a newer group [5]. The company started out moderating voice chat in online games [6]. Modulate counts more than 10 million hours of audio a month and more than 600 million in total [9]. The company did not break out how much of that volume comes from voice-agent customers. For a buyer, renewals among that group would say more than hours processed. TechCrunch reported that Modulate often runs beside whatever voice stack a company already uses, there only to analyze calls [19]. Modulate says it is working to expand on-premises and on-device deployment for privacy [21]. The thing being pitched and the thing on the price list are different products. Huffman said developers "shouldn't have to rebuild the audio intelligence layer every time they create a new voice experience" [16]. Today, a developer can buy batch transcription through the existing API at three cents an hour [17]. Huffman calls transcription the crowded part of the market: "a lot of folks are doing transcription," he told TechCrunch [18]. The accuracy evidence comes in two grades. Anyone can check the public results. The deepfake detector has ranked first on Hugging Face's Speech Deepfake Arena since March with a 1.1% equal error rate [10], and the transcription model placed first of 88 entries on the Open ASR Leaderboard in July [11]. The rest rests on Modulate's word. It says Velma is twice as accurate as general-purpose large language models at finding real problems and raises seven times fewer false alarms, comparisons The Next Web notes the company made itself [13]. It also says its ensemble design is up to 1,000 times more efficient than one large model [8]. At a hospital front desk [4], the false-alarm rate on that hospital's own lines decides whether staff keep the tool switched on. The false-alarm figure on offer is self-reported. The funding totals disagree too. Modulate says backers have put in $60 million counting this round [24]. That implies $35 million before it [1]. TechCrunch, citing PitchBook, put the pre-round figure at $41 million, at a $170 million valuation [25], or $6 million more [2]. I'd sort the decision with two questions. Is the failure in the words or in the sound? Does it have to be caught during the call or after? - Words, after the call: a wrong answer or a skipped disclosure. A transcript and the analytics a team already runs will find it. - Words, during the call: compliance rules for agents in regulated work. Modulate sells this check [20], but the rules concern what was said, so a transcript-based check is the first thing to price it against. - Sound, after the call: Huffman's polite, dissatisfied caller. Batch review of recorded agent calls can surface it without anything sitting in the live path. - Sound, during the call: a cloned voice on the line. Only this box clearly needs a separate live audio layer beside the voice stack.
What to watch
- Whether Modulate publishes prices for Velma's deepfake and agent-grading analysis once the new SDKs and APIs ship.
- Independent false-alarm results for Velma on real call-center or hospital traffic, to test the company's claim of seven times fewer false alarms.
- Delivery of on-premises and on-device deployment, the option for buyers who cannot send call audio outside their own network.