Skip to content

Build1 publisher3 min readPublished

Community MLX port makes Qwen3-TTS the fastest offline voice cloner on an M1 Pro

Qwen3-TTS went from RTF 9 under PyTorch to faster than real time through a community MLX port on one developer's M1 Pro. The port took it from written off to first place among the commercially usable cloners tested for an offline Mac app.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Community MLX port makes Qwen3-TTS the fastest offline voice cloner on an M1 Pro
Generated illustration

What happened

  • The app turns a child's reading homework into audio in a parent's cloned voice with word-by-word highlighting, and voice samples never leave the laptop.
  • Before benchmarking, F5-TTS (CC-BY-NC), XTTS-v2 (CPML) and self-hosted fish-speech (research license) were ruled out for commercial use, according to the developer.
  • Chatterbox, under the MIT license, needed 808 seconds of compute on MPS for 40 seconds of audio, an RTF of 20, and was dropped.
  • ZipVoice, a 123M-parameter Apache-2.0 model from k2-fsa weighing 468MB, led the field at near real time before Qwen3-TTS was tested.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability A Mac app can clone a parent's voice and render homework faster than real time on an M1 Pro with an Apache-2.0 model and no cloud call.
  • cost Picking Qwen3-TTS over ZipVoice makes the bundled model about five times larger, 2.4GB against 468MB, for an app that has to ship the engine inside a desktop download.
  • decision A PyTorch-on-MPS result alone had pushed this app toward a paid cloud tier, so dropping an engine for Mac now means testing its community runtime ports first.
  • exposure Shipping Qwen3-TTS on a Mac ties the app to community-maintained MLX weights and mlx-audio, because the vendor's official qwen-tts package did not run usably on MPS.

Qwen3-TTS-0.6B-Base came in with community reports of RTF 0.86 for cloning. Every one of those numbers came from CUDA [12]. On the M1 Pro, the official qwen-tts package crashed in the embedding layer under MPS with bf16. It printed "Placeholder storage has not been allocated" on both torch 2.9 and 2.13 [13]. CPU in fp32 ran for seven minutes and produced zero output [13]. That result was at least unambiguous. MPS in fp32 with device_map did run, at RTF about 9, and a ten-word sentence took roughly six minutes [14]. The CUDA reports were about ten times faster for the same weights [1].

The developer wrote the model off. The plan was a cloud API for a premium tier while watching for MPS support upstream [15]. The community had already ported the weights to MLX as mlx-community/Qwen3-TTS-12Hz-0.6B-Base-bf16, run through mlx-audio [16]. On the same machine it ran faster than real time. The developer ranked it the fastest cloning engine tested, with the best quality-to-size ratio [16]. The post's text does not include the measured RTF for that run. "If a model looks too slow on your platform, check whether someone has re-implemented it for your runtime before you write it off," the developer wrote [17].

Every figure here comes from one MacBook Pro M1 Pro with 32GB. The reference voice was a LibriVox narrator, 60 seconds of dense speech, and the test text was about 40 seconds of Peter Rabbit [5]. For the ranking to carry over to another product, the target has to be Apple Silicon running the MLX port. The reference audio also has to be clean. Qwen's in-context cloning needs a reference that is sentence-aligned at the start, with a verbatim transcript [18]. A trimmed clip fed with the full-length transcript produced output that ran to the token limit, mostly silence and noise [18]. "ZipVoice eats dirty references without blinking; Qwen does not," the developer wrote [19].

The app now prepares references the same way for every engine [20]:

1. Record the reference voice. 2. Transcribe it with Whisper. 3. Trim at sentence boundaries. 4. Feed the trimmed clip and transcript to the engine.

ZipVoice has its own sharp edges, and the post documents them well. Each per-sentence call took 10 to 16 seconds of wall clock, about 7 of them spent loading the model [8]. Loading was 44 to 70 percent of a cold call [2]. The RTF of roughly 0.6 to 1.5 holds only with the model kept resident [8]. Prompt text has to match prompt audio exactly. A 28-second clip paired with only the first sentence's transcript produced 44 seconds of that sentence repeated, while an exact 6-second clip and transcript worked [9]. Batch mode on MPS went into repetition loops on inputs that worked one sentence at a time, so the production pipeline calls it per sentence [10]. The zipvoice package on PyPI is an empty shell. The real code installs from the k2-fsa GitHub repository, along with dependencies it does not declare [11].

What to watch

  • An upstream PyTorch or qwen-tts fix for the MPS bf16 embedding-layer crash would put the official package back in the comparison.
  • A fix for ZipVoice's batch-mode repetition loops on MPS would remove the one-sentence-at-a-time limit in its production pipeline.
  • Published RTF figures for the MLX Qwen3-TTS port on other M-series chips and memory sizes would show whether the M1 Pro result generalises.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories