Matcha-TTS-PL v2 — one synthetic voice, modern Polish, English words
Second version of Matcha-TTS-PL: the same small, fast, non-autoregressive Matcha-TTS (20.9 M parameters, real-time on a small GPU or in a browser) distilled from VoxCPM2, a 2B-parameter TTS, onto a single synthetic voice. What changes compared with v1:
- One original voice,
male_voxcpm_1— designed from a text description with VoxCPM2 (OpenBMB, Apache-2.0) and frozen as a prompt cache; not a recording or a clone of a real person (closest of our reference readers: speaker similarity 0.83). - Modern Polish of a humanoid assistant — 13.1 h of training speech: assistant speech in 18 everyday domains, modern corpora (NKJP, Wikinews, KPWr, ParlaMint, OpenAssistant, Wikibooks, Wikivoyage) and reviewed YouTube transcripts; two thirds of it in multi-sentence passages (10–25 s), so prosody is planned across sentences.
- English words inside Polish sentences — "Po meetingu wyślę ci feedback na maila" is read with English phonemes:
English words are detected automatically and phonemized by espeak-ng en-us, the rest by espeak-ng pl
(
english/). No pronunciation dictionary needed for common loanwords and brands. - A HiFi-GAN fine-tuned on this voice, with real ground truth (the teacher's audio) — the v1 vocoder was tuned for the audiobook readers and blurs this voice.
Try it in the browser: spaces/machinekind/matcha-tts-pl-v2 · training data: datasets/machinekind/polish-synthetic-speech-voxcpm2.
Quality (12 held-out sentences, Mac CPU, 4 ODE steps, temperature 0.8)
| system | UTMOS ↑ | Whisper WER ↓ | similarity to male_voxcpm_1 ↑ |
|---|---|---|---|
| VoxCPM2 teacher (2B, autoregressive) | 3.68 | 1.3 % | 0.883 |
| v2 (this model + its vocoder) | 3.40 | 1.3 % | 0.850 |
| v1 model with the new voice slot, before fine-tuning | 2.91 | 2.7 % | 0.725 |
WER counts Whisper's spelling too: the teacher's two "errors" are "krutki" for krótki and "HWSR" for the robot's name.
Defaults: temperature 1.0, tempo 0.8 (length_scale; at 1.0 v2 speaks about 16 % slower than the teacher), no style token by default (8 = neutral).
Files
| path | what |
|---|---|
onnx/matcha_pl_t4.onnx, onnx/matcha_pl_t2.onnx |
acoustic model + male_voxcpm_1 vocoder in one graph, 4 / 2 ODE steps; inputs x (phoneme ids), x_lengths, scales=[temperature, length_scale], spk_emb (float32 [1, 64]); outputs wav (22.05 kHz), wav_lengths |
onnx/voices.json |
the male_voxcpm_1 speaker embedding, 18 style embeddings, defaults |
model/matcha_pl_v2.ckpt |
PyTorch checkpoint (Matcha-TTS + the tts-pl patch: style tokens, Polish cleaners) — speaker 20 = male_voxcpm_1 |
vocoder/generator_male_voxcpm_1, vocoder/config.json |
HiFi-GAN (v1 architecture) fine-tuned on this model's mels ↔ the male_voxcpm_1 audio |
english/ |
English detection (english_detect.py, english_words.txt, english_phrase_words.txt; browser: english.js, english_data.json) and the mixed phonemizer (polish_mixed_cleaners.py; browser: cleaner.js). A word that is also a Polish word (SGJP dictionary via Morfeusz 2, BSD-2; e.g. sale, to go) is read as Polish unless it is listed on its own in english_words.txt or written in braces: {sale}. Detection can be switched off: everything is then read as Polish, braces still mark English. |
samples/ |
the 12 held-out sentences, v2 and teacher |
data/ |
style token map, speaker map |
ATTRIBUTION.md |
every text source of the training data, and the base model's data |
Quick start (ONNX, Python)
import re, sys, json, numpy as np, onnxruntime as ort, soundfile as sf
sys.path.append("english") # english_detect.py + english_words.txt (optional fallback: pip install wordfreq)
from english_detect import mark_english
from matcha.text import text_to_sequence # Matcha-TTS with the tts-pl cleaners: paste english/polish_cleaners.py and
from matcha.utils.utils import intersperse # english/polish_mixed_cleaners.py into matcha/text/cleaners.py
text = mark_english("Po meetingu wyślę ci feedback na maila? A w weekend zrobię update.")
text = re.sub(r"\s+\?", "?", text) # keep '?' glued to the word, as in the training text
ids = np.array(intersperse(text_to_sequence(text, ["polish_mixed_cleaners"])[0], 0), np.int64)
vj = json.load(open("onnx/voices.json"))
spk = np.array(vj["speakers"]["0"]["emb"], np.float32) + np.array(vj["styles"]["8"]["emb"], np.float32) # 8 = neutral style (optional)
sess = ort.InferenceSession("onnx/matcha_pl_t4.onnx")
wav, n = sess.run(None, {"x": ids[None], "x_lengths": np.array([ids.size]), "scales": np.array([1.0, 0.8], np.float32), "spk_emb": spk[None]})
sf.write("out.wav", wav[0, :n[0]], 22050) # scales = [temperature, length_scale (tempo)]
Text front end: mark_english() wraps English words in braces (Po {meeting}u wyślę ci {feedback}.), the cleaner
phonemizes braced spans with espeak-ng en-us and the rest with pl (with the Polish ending of an inflected English
word glued on unstressed, and one-letter prepositions before English words handled by rule). You can brace any word
yourself. Questions: keep the ? glued to the word ("kuchni?"), as in the training text — the model follows
VoxCPM2's question intonation (wh-questions fall, yes/no questions rise); a space before ? makes wh-questions rise. The browser port produces the training phonemes character for character (tested on 300 sentences).
Training
- Base: Matcha-TTS-PL v1 (target model), new speaker row initialised from the closest real reader; same mel statistics, style tokens and phoneme table (all en-us phones were already in the symbol table, now trained).
- Data: machinekind/polish-synthetic-speech-voxcpm2 — 6 493 clips, 13.1 h, generated by VoxCPM2 with the frozen male_voxcpm_1 voice (anchor + continuation, 10 diffusion steps), every clip gated by Whisper (text match), UTMOS ≥ 3.2 and speaker similarity ≥ 0.80, up to 3 takes per sentence. Source texts reviewed and minimally completed by Claude (cut sentences finished, numbers written as words).
- Matcha fine-tune: 10 k steps, batch 64, lr 5e-5, bf16, RTX 6000 Ada (≈ 1.5 h); quality plateaued from 5 k steps.
- Vocoder: teacher-forced mels of the 10 k checkpoint for all training clips paired with the VoxCPM2 audio, HiFi-GAN universal v1 fine-tuned 30 k steps, lr 2e-5.
Limitations
- One voice. The rhythm is more regular than the teacher's: Matcha's deterministic duration predictor averages timing.
- English detection uses a curated list (≈ 1 900 words and brands, Polish lookalikes excluded); an unlisted English word is read the Polish way unless you put it in braces.
- Synthetic voice: label generated speech as synthetic; don't present it as a real person.
Licence
CC BY-SA 4.0 (training data includes CC BY-SA texts; the base model is CC BY-SA 4.0). VoxCPM2 is Apache-2.0;
espeak-ng (runtime phonemizer) is GPL-3.0. See ATTRIBUTION.md.
Model tree for machinekind/Matcha-TTS-PL-v2
Base model
machinekind/Matcha-TTS-PL