Maintained by hexgrad · Apache 2.0 · v1.0, voice af_heart, seed 0 · One request: 45,600 samples at 24 kHz (about 1.9 seconds), medians of 60 runs · Bulk: 1,000 LibriSpeech sentences, median of 5 passes
| Metric | NukeTorch | MLX | PyTorch | vs MLX | vs PyTorch |
|---|---|---|---|---|---|
| One request, text to finished audio (ms) | 18.47 | 46.65 | 61.68 | 2.53× | 3.34× |
| 1,000 sentences in bulk (s) § | 46.36 | 156.38 | not run | 3.37× | n/a |
| Real-time factor, one request (×) | 102.85 | 40.73 | 30.80 | 2.53× | 3.34× |
| Real-time factor, bulk (×) | 150.42 | 44.57 | not run | 3.37× | n/a |
| Transcription error rate, bulk † | 2.25% | 2.12% | not run | 0.13 points more | n/a |
| Peak memory, one request (MB) | 536 | 2,048 | 2,150 | 74% less | 75% less |
| Peak memory, 1,000 sentences (MB) § | 2,355 | 7,987 | not run | 71% less | n/a |
§ NukeTorch in its default bulk mode, which schedules the requests itself; MLX generating them one after another, as it ships. Median of five passes; the slowest NukeTorch pass was 46.38 s and the fastest MLX pass 156.27 s. PyTorch was not run in bulk. Memory is the process’s physical footprint sampled from outside, the same probe for every engine, in runs separate from the timings; MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set. † The difference comes from NukeTorch’s own text-to-phoneme step, not the synthesis: fed the same Misaki phonemes as MLX, NukeTorch scores 2.14% against 2.12%.
What happens technically
Kokoro is small enough that very little of a request is arithmetic. On the test clip a single request finishes after 18.5 ms instead of 46.7 ms under MLX. The work being done is identical; what is gone is the routing around each operation. In bulk the gap widens to 3.4×, because NukeTorch also schedules the requests itself while MLX works through them one at a time. The five bulk passes landed within a third of a second of each other, 46.1 to 46.4 seconds, so this is a steady result rather than a lucky run. The memory result has the same cause. A general framework keeps a graph, an allocator and its own intermediate buffers alive for every operation it might be asked to run next, and none of that is needed when the sequence of operations is fixed in advance.
What it means for you
On a single sentence on a fast Mac, the time saved is small in absolute terms. It adds up when Kokoro is one step in a chain. In a local voice assistant, transcription, the language model and speech each add a wait, and about 28 ms back from the last step is 28 ms nobody sits through. For bulk narration, 1,000 sentences take 46 seconds instead of 2.6 minutes.
The memory matters in every setup. About 0.54 GB instead of 2 GB is room for a larger language model on the same Mac, or a lower minimum spec for an app that ships Kokoro. If you maintain Kokoro, it is also a better answer for your readme's Mac section than PyTorch's Apple fallback.
Does the output still match
Qualified: 42 requests across four languages. Waveform cosine similarity of at least 0.999998 against the reference, where the gate was set at 0.999. Relative RMSE at most 0.20%, against a 5% gate. Identical sample counts. Word error rate from automatic transcription equal on all 36 English requests. Deliberately broken control cases fail the same gates, so the gates are doing work.
On the 1,000-sentence bulk test, NukeTorch’s output has 2.25% of words wrong against MLX’s 2.12%, 0.13 percentage points more. The cause is NukeTorch’s own text-to-phoneme step, which follows Misaki’s rules without its part-of-speech tagger and without its fallback for words outside the dictionary. Given the same Misaki phonemes, NukeTorch scores 2.14%, level with MLX.
To answer the obvious objection, that we beat slow settings, NukeTorch was also run exactly the way MLX and PyTorch run: FP32 weights, the same Misaki text-to-phoneme step, and one request at a time in bulk. It is still 2.35× faster than MLX on a single request (19.88 ms against 46.65 ms), 3.10× faster than PyTorch, and 2.58× faster than MLX across the 1,000 sentences (60.6 s against 156.4 s). Transcription error rates match: 2.14% against 2.12%. Peak memory on a single request is 636 MB against MLX’s 2,048 MB. Most of the speed is the runtime itself; half precision and NukeTorch’s own bulk scheduling add the rest.
Nuance
- The single-request clip is short: about 1.9 seconds of audio from a three-word sentence. Short clips flatter a runtime with low fixed overhead, which is exactly what we claim to be. The bulk test uses 1,000 real sentences instead; passages of several minutes are not yet measured.
- The bulk comparison is NukeTorch's default bulk mode against MLX generating one request at a time, as it ships. It compares the two as you would run them, not identical scheduling.
- All three engines were measured on the same machine. Single requests are medians of 60 runs, timed from raw text to finished audio on every engine; the slowest NukeTorch run was 18.80 ms and the fastest MLX run 46.03 ms. No P95 or P99 claim, here or anywhere on this site.
- Precision: NukeTorch runs FP16 with FP32 accumulation. PyTorch and MLX run FP32, their default configuration, as their packages ship. We did not tune them. NukeTorch's default is qualified on output rather than on bit-for-bit parity. Run at FP32 with MLX’s own phonemes, it is still 2.35× faster per request and 2.58× in bulk.
- Bulk transcription error rate is 0.13 points higher than MLX’s because of NukeTorch’s own text-to-phoneme step; with Misaki’s phonemes the rates match.
How to use it
macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined