Invest1 publisher3 min readPublished
Modulate raises $25 million to sell the tone and deepfake signals a 3-cent transcript misses
Modulate raised $25 million led by Future Ventures, lifting the Boston audio-AI company to $60 million in total funding. With transcripts priced at 3 cents an hour, its revenue case rests on the tone and deepfake signals sold on top.
The Investor · Invest desk

What happened
- Modulate says its models analyze more than 10 million hours of audio a month and have processed more than 600 million hours in total.
- Its Velma platform runs on an ensemble of more than 100 specialized audio models, which Modulate says is up to 1,000 times more efficient than one large model.
- Modulate reports that its deepfake detection reached 98.9% accuracy on public benchmark data.
- The new money goes to research, engineering, developer relations and partnerships, plus industry models, SDKs and APIs for outside developers.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- cost At transcript rates Modulate's monthly volume bills about $3.6 million a year against $60 million raised, so the investors' return depends on tone and fraud signals selling well above 3 cents an hour.
- decision Banks, hospitals and contact centers are being offered Modulate's deepfake screening as a layer inside other vendors' security products and voice platforms, so those vendors decide whether it reaches them.
- constraint Modulate's claims of twice the true positives and seven times fewer false positives than traditional LLMs are self-reported, so a fraud team still has to measure false alarms on its own calls.
Batch transcription through Modulate's API costs $0.03 an hour [5]. Put all 10 million of its monthly hours through that meter [4] and the bill comes to $300,000 a month, or about $3.6 million a year [2]. That sits against $60 million of capital raised to date, $35 million of it before this round [2][1]. The company did not disclose revenue or a price for anything above a transcript, so the $300,000 is what the volume would bill at transcript rates. It is not an estimate of sales.
The 600 million hours processed to date equal 60 months at the current pace [4][3]. A company whose volume had just surged would show a far smaller multiple.
The case for everything above the transcript is that a transcript throws information away. "Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can't be solved from a transcript," Modulate said in its announcement [15]. Velma reads emotion, tone, intent and synthetic speech from the audio itself [3] and combines them into events such as suspected fraud, harassment or problems with a voice AI agent [10].
Whether this round shows deepfake detection becoming a funded category of its own depends on the spending plan. The plan points at developers [11], and the stated goal is a layer that other companies build into voice agents, security products and communications platforms without making their own audio models [12]. Fraud and deepfake detection sit on the use list beside agent supervision, customer experience and trust and safety [9], with harassment and child-grooming detection and voice masking further down [17]. Future Ventures, Hyperplane and Lakestar [1] are paying for a component that goes inside someone else's fraud product.
The round supports three readings. Deepfake detection could be the paid product, with cheap transcripts bringing in the hours. Agent supervision could be the bigger line, because businesses running voice agents need to check whether those agents understand customers and respond as intended [14]. Or the margin could come from the ensemble design [8], so that low per-hour pricing across every use is the business itself.
I think the second reading fits the spending plan best, and fraud is simply the use case easiest to explain to an investor. The counter-case is that financial services firms, healthcare groups and contact centers, the buyers Modulate names for screening suspicious calls [13], will pay more per flagged call than anyone pays to supervise a voice agent. On that view fraud carries the revenue while being one line among several. My view is wrong if fraud and synthetic-speech screening turn out to account for most of those 10 million monthly hours [4].
What to watch
- A published per-hour price for Velma's deepfake or fraud signals, which would show how far above the 3-cent transcript rate the paid product sits.
- A named bank, hospital or contact-center deployment reporting false-positive rates on live calls to set beside the 98.9% benchmark figure.
- A company selling only deepfake or voice-fraud detection raising on its own, which would test whether investors fund the use case apart from a broader audio layer.