NVIDIA · CC BY 4.0 · 25 languages · One clip: 6.4 seconds of English speech, median of 60 calls · Bulk: 1,000 LibriSpeech recordings, 7,389 seconds of speech, median of 5 passes
| Metric | NukeTorch | MLX | PyTorch | vs MLX | vs PyTorch |
|---|---|---|---|---|---|
| One clip, request time (ms) | 20.09 | 46.90 | 85.79 | 2.33× | 4.27× |
| 1,000 clips (s) § | 14.14 | 48.05 | not run | 3.40× | n/a |
| Clips per second, bulk | 70.72 | 20.81 | not run | 3.40× | n/a |
| Real-time factor, one clip (×) | 318.6 | 136.5 | 74.6 | 2.33× | 4.27× |
| Real-time factor, bulk (×) | 522.6 | 153.8 | not run | 3.40× | n/a |
| Word error rate, 1,000 clips | 1.77% | 1.75% | not run | about the same | n/a |
| Peak memory, one clip (MB) | 1,741 | 3,174 | 3,891 | 45% less | 55% less |
| Peak memory, 1,000 clips (MB) § | 2,253 | 45,773 | not run | 95% less | n/a |
| Model file on disk (MB) | 1,196 | 2,392 | 2,392 | half the size | half the size |
| Model inference, one clip (ms) ‡ | 19.48 | 42.14 | 79.92 | 2.16× | 4.10× |
§ 1,000 LibriSpeech test-clean recordings, 7,389 seconds of speech, median of five passes per engine. NukeTorch runs in its default bulk mode, which schedules the clips itself; MLX transcribes them one after another, as it ships. PyTorch was not run in bulk. Memory is the process’s physical footprint sampled from outside, the same probe for every engine, in runs separate from the timings. MLX keeps memory it has finished with in its own cache, so its bulk peak reflects its largest working set. ‡ A diagnostic, not a like-for-like comparison: NukeTorch is timed with the graphics chip’s own clock, MLX and PyTorch with the computer’s main clock. One-clip request time is the closer comparison.
What happens technically
Parakeet's decoder works through the audio step by step, emitting up to ten symbols per slice of sound, and each step is a handful of small operations that each have to be routed before they run. That is the shape of work NukeTorch is built for: 2.3× faster than MLX on a single clip, and 3.4× across 1,000 recordings, where NukeTorch also schedules the clips itself. Memory is the other half of the result. Over the bulk job MLX’s footprint climbs to 45.8 GB, partly because it holds on to memory it has finished with; NukeTorch stays at 2.3 GB, and stores the model at half the size on disk.
What it means for you
Two hours of recorded speech, 1,000 clips, transcribed in 14 seconds instead of 48 under MLX, with the same word error rate, in 2.3 GB of memory instead of 45.8 GB. For meeting notes, call recordings or podcast archives, that is a batch job that fits on an ordinary Mac.
In a local voice assistant, listening is the first wait in every turn. A 6.4-second utterance becomes text in 20 ms instead of 47 ms, before the language model and the voice even start.
Nuance
- Precision: NukeTorch runs FP16. MLX and PyTorch use the FP32 weights as shipped, with MLX computing the audio features in BF16. The word error rates match, but it is not a same-precision comparison.
- The bulk comparison is NukeTorch’s default bulk mode against MLX one clip at a time, as it ships.
- MLX keeps memory it has finished with in its own cache, so its 45.8 GB bulk peak reflects its largest working set rather than a steady need. Its median over the second half of the run was still 41.5 GB, against NukeTorch’s 2.3 GB.
- The one-clip sample was generated with Kokoro, not recorded. The 1,000-clip test uses real recordings.
- English only, although Parakeet v3 supports 25 languages.
How to use it
macOS on Apple silicon · the same model files and checkpoint revision you use today · licensing undetermined