Instructions to use mudler/parakeet-cpp-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use mudler/parakeet-cpp-gguf with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("mudler/parakeet-cpp-gguf") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
Parakeet GGUF โ models for parakeet.cpp
GGUF-format weights for parakeet.cpp, a C++/ggml port of NVIDIA NeMo Parakeet that matches the upstream PyTorch models on CPU. This single repo collects every supported model ร quantization as a flat set of .gguf files โ download just the one you need.
F16 is the recommended default โ same accuracy as F32, ~1.7ร smaller, and typically the fastest on modern CPUs via ggml's F32รF16 matmul fast path.
Models
tdt_ctc-110m
Source: nvidia/parakeet-tdt_ctc-110m ยท Hybrid TDT+CTC (FastConformer) ยท heads: TDT + CTC
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt_ctc-110m-f16.gguf โ recommended |
F16 | 267.5 MB | 0.0000 |
tdt_ctc-110m-q8_0.gguf |
Q8_0 | 177.8 MB | 0.0000 |
tdt_ctc-110m-q6_k.gguf |
Q6_K | 155.9 MB | not measured |
tdt_ctc-110m-q5_k.gguf |
Q5_K | 143.3 MB | not measured |
tdt_ctc-110m-q4_k.gguf |
Q4_K | 131.4 MB | 0.0000 |
realtime_eou_120m-v1
Source: nvidia/parakeet_realtime_eou_120m-v1 ยท Cache-aware streaming RNNT (FastConformer, EOU/EOB) ยท heads: RNNT (streaming) ยท License: NVIDIA Open Model License
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
realtime_eou_120m-v1-f16.gguf โ recommended |
F16 | 266.5 MB | not measured |
realtime_eou_120m-v1-q8_0.gguf |
Q8_0 | 176.0 MB | not measured |
realtime_eou_120m-v1-q6_k.gguf |
Q6_K | 153.9 MB | not measured |
realtime_eou_120m-v1-q5_k.gguf |
Q5_K | 141.2 MB | not measured |
realtime_eou_120m-v1-q4_k.gguf |
Q4_K | 129.1 MB | not measured |
ctc-0.6b
Source: nvidia/parakeet-ctc-0.6b ยท CTC (FastConformer) ยท heads: CTC
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
ctc-0.6b-f16.gguf โ recommended |
F16 | 1373.4 MB | 0.0000 |
ctc-0.6b-q8_0.gguf |
Q8_0 | 875.4 MB | 0.0000 |
ctc-0.6b-q6_k.gguf |
Q6_K | 746.8 MB | not measured |
ctc-0.6b-q5_k.gguf |
Q5_K | 676.3 MB | not measured |
ctc-0.6b-q4_k.gguf |
Q4_K | 609.9 MB | not measured |
rnnt-0.6b
Source: nvidia/parakeet-rnnt-0.6b ยท RNNT transducer (FastConformer) ยท heads: RNNT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
rnnt-0.6b-f16.gguf โ recommended |
F16 | 1402.8 MB | 0.0000 |
rnnt-0.6b-q8_0.gguf |
Q8_0 | 903.9 MB | 0.0000 |
rnnt-0.6b-q6_k.gguf |
Q6_K | 776.3 MB | not measured |
rnnt-0.6b-q5_k.gguf |
Q5_K | 705.7 MB | not measured |
rnnt-0.6b-q4_k.gguf |
Q4_K | 639.2 MB | not measured |
tdt-0.6b-v2
Source: nvidia/parakeet-tdt-0.6b-v2 ยท TDT transducer (FastConformer) ยท heads: TDT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt-0.6b-v2-f16.gguf โ recommended |
F16 | 1404.2 MB | 0.0000 |
tdt-0.6b-v2-q8_0.gguf |
Q8_0 | 903.8 MB | 0.0000 |
tdt-0.6b-v2-q6_k.gguf |
Q6_K | 775.9 MB | not measured |
tdt-0.6b-v2-q5_k.gguf |
Q5_K | 705.0 MB | not measured |
tdt-0.6b-v2-q4_k.gguf |
Q4_K | 638.4 MB | not measured |
tdt-0.6b-v3
Source: nvidia/parakeet-tdt-0.6b-v3 ยท TDT transducer (FastConformer) ยท heads: TDT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt-0.6b-v3-f16.gguf โ recommended |
F16 | 1441.0 MB | 0.0000 |
tdt-0.6b-v3-q8_0.gguf |
Q8_0 | 940.7 MB | 0.0000 |
tdt-0.6b-v3-q6_k.gguf |
Q6_K | 812.7 MB | not measured |
tdt-0.6b-v3-q5_k.gguf |
Q5_K | 741.9 MB | not measured |
tdt-0.6b-v3-q4_k.gguf |
Q4_K | 675.2 MB | not measured |
ctc-1.1b
Source: nvidia/parakeet-ctc-1.1b ยท CTC (FastConformer) ยท heads: CTC
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
ctc-1.1b-f16.gguf โ recommended |
F16 | 2395.8 MB | 0.0000 |
ctc-1.1b-q8_0.gguf |
Q8_0 | 1526.3 MB | 0.0000 |
ctc-1.1b-q6_k.gguf |
Q6_K | 1301.7 MB | not measured |
ctc-1.1b-q5_k.gguf |
Q5_K | 1178.5 MB | not measured |
ctc-1.1b-q4_k.gguf |
Q4_K | 1062.6 MB | not measured |
rnnt-1.1b
Source: nvidia/parakeet-rnnt-1.1b ยท RNNT transducer (FastConformer) ยท heads: RNNT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
rnnt-1.1b-f16.gguf โ recommended |
F16 | 2425.2 MB | 0.0000 |
rnnt-1.1b-q8_0.gguf |
Q8_0 | 1554.7 MB | 0.0000 |
rnnt-1.1b-q6_k.gguf |
Q6_K | 1331.2 MB | not measured |
rnnt-1.1b-q5_k.gguf |
Q5_K | 1207.9 MB | not measured |
rnnt-1.1b-q4_k.gguf |
Q4_K | 1091.9 MB | not measured |
tdt-1.1b
Source: nvidia/parakeet-tdt-1.1b ยท TDT transducer (FastConformer) ยท heads: TDT
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt-1.1b-f16.gguf โ recommended |
F16 | 2425.3 MB | 0.0000 |
tdt-1.1b-q8_0.gguf |
Q8_0 | 1554.8 MB | 0.0000 |
tdt-1.1b-q6_k.gguf |
Q6_K | 1331.2 MB | not measured |
tdt-1.1b-q5_k.gguf |
Q5_K | 1207.9 MB | not measured |
tdt-1.1b-q4_k.gguf |
Q4_K | 1091.9 MB | not measured |
tdt_ctc-1.1b
Source: nvidia/parakeet-tdt_ctc-1.1b ยท Hybrid TDT+CTC (FastConformer) ยท heads: TDT + CTC
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
tdt_ctc-1.1b-f16.gguf โ recommended |
F16 | 2429.5 MB | 0.0000 |
tdt_ctc-1.1b-q8_0.gguf |
Q8_0 | 1559.0 MB | 0.0000 |
tdt_ctc-1.1b-q6_k.gguf |
Q6_K | 1335.4 MB | not measured |
tdt_ctc-1.1b-q5_k.gguf |
Q5_K | 1212.1 MB | not measured |
tdt_ctc-1.1b-q4_k.gguf |
Q4_K | 1096.1 MB | not measured |
WER (word error rate) is computed against the upstream NeMo reference on
tests/fixtures/speech.wav(LibriSpeech2086-149220-0033, ~7.4 s, English). 0.0 = byte-for-byte identical transcript. See parity.md and quantization.md.
nemotron-3.5-asr-streaming-0.6b
Source: nvidia/nemotron-3.5-asr-streaming-0.6b ยท Cache-aware streaming RNNT (FastConformer, 24 encoder layers), multilingual (40 language locales), conditioned on a language prompt ยท heads: RNNT (streaming, prompt-conditioned) ยท License: OpenMDW 1.1
| File | Variant | Size | WER vs NeMo |
|---|---|---|---|
nemotron-3.5-asr-streaming-0.6b-f16.gguf โ recommended |
F16 | 1484.3 MB | 0.0000 |
nemotron-3.5-asr-streaming-0.6b-q8_0.gguf |
Q8_0 | 983.7 MB | 0.0000 |
nemotron-3.5-asr-streaming-0.6b-q6_k.gguf |
Q6_K | 855.7 MB | 0.0000 |
nemotron-3.5-asr-streaming-0.6b-q5_k.gguf |
Q5_K | 784.8 MB | 0.0000 |
nemotron-3.5-asr-streaming-0.6b-q4_k.gguf |
Q4_K | 718.1 MB | 0.0000 |
This model runs both offline and as a cache-aware streaming model, and it is the only one here that takes a target language. Pass
--lang <locale>(for exampleen-US,de-DE,es-ES,ja-JP), or leave it at the defaultautoto let the model detect the language. The WER column is measured against NeMo, offline, on the languages en, de and auto; streaming and offline output also match NeMo for en, de, es, ja-JP and auto. See parity.md. The small prompt layers, the LSTM and the feature extractor stay F32 in every quantization.
huggingface-cli download mudler/parakeet-cpp-gguf nemotron-3.5-asr-streaming-0.6b-f16.gguf --local-dir models/
build/examples/cli/parakeet-cli transcribe --model models/nemotron-3.5-asr-streaming-0.6b-f16.gguf --input audio.wav --lang de-DE
nemotron-3-diarization
Source: nvidia/Nemotron-3-Diarization ยท Speaker diarization (Sortformer, up to 8 speakers, streaming speaker cache) ยท License: OpenMDW 1.1
| File | Variant | Size | Segments vs NeMo |
|---|---|---|---|
nemotron-3-diarization-f16.gguf โ recommended |
F16 | 200.7 MB | identical |
nemotron-3-diarization-q8_0.gguf |
Q8_0 | 108.7 MB | identical on 2 of 3 clips, 99.5% of frames on the third |
Checked against NeMo on a 23.6 s and a 68.5 s two-speaker clip (offline and streaming: every segment identical, to the 10 ms frame) and a 12.3 min three-speaker clip (F16 100%, Q8_0 99.9% of speech frames). Q8_0 can flip a frame whose probability sits right at the 0.5 threshold, splitting a segment (99.5% of frames on a 31.5 s clip); use F16 when segment-exact output matters. Answers "who spoke when"; pair it with any ASR model above for speaker-attributed transcripts. See diarization.md.
build/examples/cli/diarize models/nemotron-3-diarization-f16.gguf meeting.wav
ultra and redux (Moondream)
Source: moondream/parakeet-ultra and moondream/parakeet-redux by Moondream, derived from NVIDIA's nvidia/parakeet-tdt-0.6b-v3 ยท TDT transducer (FastConformer), multilingual, with a voice-activity head (transcribe --vad) ยท heads: TDT ยท license: CC-BY-4.0
| File | Variant | Size | Runs on |
|---|---|---|---|
ultra-f16.gguf |
F16 | 1441.9 MB | any backend |
ultra-q8_0.gguf |
Q8_0 | 941.5 MB | any backend |
redux-packed.gguf |
packed ternary encoder | 213.3 MB | CPU only, offline only |
redux-f16.gguf |
F16, dequantized | 1441.9 MB | any backend |
redux-q8_0.gguf |
Q8_0, dequantized | 941.5 MB | any backend |
There is no NeMo reference for these models, so there is no WER column. Each file was checked to decode the
tests/fixtures/speech.wavclip to the expected sentence. See ternary.md for the measurements.
- Ultra has ordinary F16 weights and runs on any backend, like the other v3 files.
- Redux packed keeps the encoder linear layers as ternary weights (-1, 0 or +1 times a per-group scale) and runs a native CPU kernel. It does not run on GPU backends or in streaming mode, and parakeet.cpp refuses to load it there. For GPU use, take
redux-f16.gguforredux-q8_0.gguf. - Redux F16 and Q8_0 are dequantized: the ternary weights were expanded to ordinary weights and then stored as F16 or Q8_0. They are not packed.
- Changes: these files are converted here, not trained. Nothing was trained or fine-tuned by the parakeet.cpp project. The models were trained by NVIDIA (the base) and Moondream (Ultra and Redux).
huggingface-cli download mudler/parakeet-cpp-gguf redux-packed.gguf --local-dir models/
build/examples/cli/parakeet-cli transcribe --model models/redux-packed.gguf --input audio.wav
# Long audio: cut at pauses with the model's voice-activity head (offline only).
build/examples/cli/parakeet-cli transcribe --model models/ultra-q8_0.gguf --input long.wav --vad
silero-vad (Silero Team)
Source: snakers4/silero-vad v6.2.3 by the Silero Team ยท a small voice-activity detector, not a speech recognizer ยท 8 kHz and 16 kHz in one file ยท license: MIT
| File | Variant | Size |
|---|---|---|
silero-vad-f32.gguf |
F32 | 2.2 MB |
silero-vad-f16.gguf |
F16 weights, widened to F32 at load | 1.3 MB |
- The files are converted from the official ONNX model with
scripts/convert_silero_vad_to_gguf.pyin parakeet.cpp. The GGUF records the source version and the ONNX sha256. They were converted here, not trained. - They need a parakeet.cpp build that has the standalone VAD API. That code is not in a release yet, so check the parakeet.cpp repository before relying on it.
- Use it as a stand-alone detector (
parakeet-cli vad --model silero-vad-f16.gguf --input audio.wav) or to cut long audio before transcribing with any model, including those that have no VAD head (parakeet-cli transcribe --model tdt-0.6b-v3-q8_0.gguf --input long.wav --vad --vad-model silero-vad-f16.gguf). - Probabilities match onnxruntime to about 1e-6 (F32) and 1e-3 (F16) on the test clips.
vad heads (Moondream)
Source: the voice-activity head, the mel front end and the subsampler of moondream/parakeet-redux and moondream/parakeet-ultra by Moondream, derived from NVIDIA's nvidia/parakeet-tdt-0.6b-v3 ยท VAD only: these files cannot transcribe ยท license: CC-BY-4.0
| File | Parent | Size |
|---|---|---|
redux-vad.gguf |
redux (the head and subsampler are not ternary) | 9.9 MB |
ultra-vad-q8_0.gguf |
ultra-q8_0 | 6.0 MB |
- Each file holds 20 tensors copied byte for byte from its parent, with no requantization, and records the parent file, its size and its sha256 in the metadata. They were cut out here, not trained.
- Output is identical to the full parent file:
parakeet-cli vad --probabilitiesgave byte-identical JSON on a speech clip, a noisy clip and a 600 s talk. - Against the 213 MB to 941 MB parents, load time falls from about 0.1 to 0.4 s to a few milliseconds, and memory for a 33 s clip from 0.6 to 1.1 GiB to about 0.25 GiB. Speed is the same as the parent's head.
- They need a parakeet.cpp build that can load a VAD-only file. That code is not in a release yet, so check the parakeet.cpp repository before relying on it. Use
parakeet-cli vad --model redux-vad.gguf --input audio.wav. - For a stand-alone detector Silero (above) is smaller and faster per core. The head is useful when you want its recall or already work with the Moondream models.
Quantization notes
Quantization is applied only to the large linear weights fed directly into ggml_mul_mat (encoder FFN + attention projections, subsampling output projection, joint enc/pred projections). All other tensors (mel filterbank, LSTM prediction net, conv kernels, batch_norm stats, norms, biases, embeddings) stay F32.
Usage
# 1. Clone + build parakeet.cpp
git clone https://github.com/mudler/parakeet.cpp
cd parakeet.cpp
cmake -B build -DPARAKEET_BUILD_CLI=ON && cmake --build build -j
# 2. Download one quant (F16 recommended)
huggingface-cli download mudler/parakeet-cpp-gguf tdt_ctc-110m-f16.gguf --local-dir models/
# 3. Transcribe
build/examples/cli/parakeet-cli transcribe \
--model models/tdt_ctc-110m-f16.gguf \
--input audio.wav
License
Licences differ by model, so the front matter says license: other. Each file family follows the licence of the model it was converted from:
tdt_ctc-110m-*,ctc-*,rnnt-*,tdt-*andtdt_ctc-1.1b-*: derived from NVIDIA NeMo Parakeet checkpoints released under CC-BY-4.0.realtime_eou_120m-v1-*: derived from nvidia/parakeet_realtime_eou_120m-v1, governed by the NVIDIA Open Model License.nemotron-3.5-asr-streaming-0.6b-*andnemotron-3-diarization-*: derived from NVIDIA Nemotron models, governed by the OpenMDW License Agreement, version 1.1.ultra-*andredux-*: see below.
The notes below add detail. ultra-*.gguf and redux-*.gguf are converted from moondream/parakeet-ultra and moondream/parakeet-redux by Moondream, which are derived from NVIDIA's parakeet-tdt-0.6b-v3. Both are also CC-BY-4.0: credit Moondream and NVIDIA when you use these files. They were converted here, not trained, and the Redux F16 and Q8_0 files are dequantized from the ternary weights. redux-vad.gguf and ultra-vad-q8_0.gguf hold only the VAD head, front end and subsampler of the Moondream models, cut out of the files above under the same CC-BY-4.0 terms: credit Moondream and NVIDIA; they were cut out here, not trained. silero-vad-*.gguf is converted from Silero VAD v6.2.3 and is released under the MIT license, Copyright (c) 2020-present Silero Team; it was converted here, not trained. nemotron-3-diarization-*.gguf is derived from nvidia/Nemotron-3-Diarization and nemotron-3.5-asr-streaming-0.6b-*.gguf from nvidia/nemotron-3.5-asr-streaming-0.6b. Both are governed by the OpenMDW License Agreement, version 1.1. The parakeet.cpp runtime is MIT-licensed.
- Downloads last month
- 24,151
Model tree for mudler/parakeet-cpp-gguf
Base model
moondream/parakeet-redux