Compare the deployment boundaries
| Candidate | Deployment | License boundary | Primary source |
|---|---|---|---|
| Kokoro-82M | Local 82M model with preset voices; verify the available voice for your language. | Apache-2.0 model weights; Kokoro v1.0 does not ship a Polish voice. | Model card ↗ |
| Piper (OHF engine) | Local ONNX synthesis, CLI/server/library paths and downloadable voices. | Maintained engine: GPL-3.0. Downloaded voice terms are specific to the voice; the repository license does not replace them. | Maintained engine ↗ |
The archived Rhasspy engine used MIT; that does not describe the maintained engine. Neither candidate is an arbitrary voice-cloning default.
Evidence status
The May hard-text rows are withheld: their manifests contain placeholder hashes and do not account for the claimed 30 audio samples. No auditable measured hard-text ranking is available until complete artifacts are restored.
Previous Kokoro winner claims based on those rows are withdrawn. A separate preference study cannot replace a matched Piper-versus-Kokoro test using the same device, prompts and voices.
Run a useful comparison
- Select the exact voice and checkpoint for each system. Check your language and license requirements before benchmarking.
- Use identical clean sentences, names, numbers, addresses and long paragraphs. Keep volume and audio format comparable.
- Measure model loading, warm first audio, full synthesis time, peak memory and error counts separately on your target CPU.
- Store audio and transcripts with real content hashes. Run blind listener comparisons and inspect pronunciation failures.