Key takeaways
- Use language, intelligibility, and timing metrics to shortlist candidates; audition performance in the actual scene.
- Measure first audible playback and interruptions on your target device, separately from model-only timing.
- Keep model access terms, the exact voice-file license, attribution, and performer consent in separate records.
Use the Open TTS Leaderboard to narrow your NPC voice shortlist, then audition the finalists with dialogue from the kind of scene you are building. A clear transcript, a quick first sound, and convincing acting answer different questions. Start with the required language and deployment target; choose the voice only after checking its performance and permissions.

What the new leaderboard can help you decide
Hugging Face introduced the Open TTS Leaderboard on September 30, 2026. Its default ranking uses English results, with multilingual and voice-cloning views available. The authors caution that intelligibility and speaker-similarity scores do not directly measure naturalness, expressiveness, or listener preference. A high rank therefore does not establish that a voice can sell your villain’s threat or your merchant’s joke.[1]
First write a one-sentence job description: “An English merchant delivers short, prewritten quest lines,” or “A companion reads newly generated responses on a laptop.” Those are different acceptance problems. For baked dialogue, prioritize usable takes and consistent delivery. For live responses, add startup delay, sustained playback, and cancellation to the same audition.
Read each metric as one part of an audition
| Signal | What it describes | What to check in your scene |
|---|---|---|
| WER / CER | Transcript error at word or character level; lower is better | Names, quantities, negation, and subtitle agreement |
| RTFx | Audio duration divided by generation time; higher is faster | Whether the chosen device keeps up while the game runs |
| TTFA | Delay until the first playable audio is available | Time from the game’s request to actual audible playback |
| Speaker similarity | Embedding similarity to reference audio in cloning evaluation | Character consistency across moods, with authorized voice material |
Match the language and test mode before comparing rows. The leaderboard’s batched offline throughput and single-request streaming results serve different purposes. Its H200 GPU and CPU results are separate hardware settings. Treat the Listen tab as a place to inspect example outputs, then build your own short scene audition.[2]
HF’s streaming protocol uses batch size one and 50 English prompts, discarding three warm-up runs and reporting median time-to-first-audio. A non-streaming model must finish the utterance before playback can begin. These conditions explain what the published timing covers; they do not measure your game’s full response path.[1]
Measure the delay the player actually experiences
Put timestamps at four boundaries: the game requests speech; synthesis starts; a playable chunk reaches the client; and sound reaches the playback device. Also record when text became available if another system produces the line. Report both the model interval and the request-to-audible interval. Keep cold starts separate from an already loaded session.
Repeat a short acknowledgment and a long instruction while your scene is active. Record a distribution rather than just the fastest attempt. Then interrupt the long line halfway through: does queued audio stop, does the next line start cleanly, and do subtitles advance correctly? A fast first chunk should not hide a stall later in the sentence.
- Set a maximum acceptable delay for the specific interaction before testing. A warning bark and a leisurely quest explanation need separate budgets.
- Keep the device, game build, model revision, voice file, language, and generation settings fixed when comparing candidates.
- Log failures as well as successful takes. A missing word or repeated line can matter more than a small average speed gain.
Pocket TTS is one candidate, not a universal winner
Kyutai describes Pocket TTS as a small, CPU-oriented streaming system with multilingual support. Its card reports roughly 200 ms to first audio and about six-times real-time generation on a MacBook Air M4 using two CPU cores. Those are Kyutai’s reported conditions, not measurements from this guide or a guarantee for a player’s device.[3]
A CPU-oriented candidate is useful to evaluate when your project needs local speech, but choose based on your actual workload. Begin with one authorized voice and one language. Do not extrapolate a short, clean example to every accent, invented name, emotional direction, or busy scene. No Pocket TTS audio was generated or listened to for this article.
Use a ten-line audition before committing a whole script
The downloadable JSON worksheet pairs the source-based decision matrix with ten newly written lines. Every result field is blank. The lines deliberately exercise different failure cases; a direction such as “whisper” is an audition goal, not a promise that every model supports a matching control.
| Test | Line | Listen or inspect for |
|---|---|---|
| Names | Bring the lantern to Mirewatch before dawn. | Stable pronunciation of the place name |
| Quantity | Three brass keys. Only one opens the north gate. | Correct number and restrictive meaning |
| Urgency | Wait. That bridge is moving. | Urgency without losing the last word |
| Quiet | Keep your voice down; the walls are listening. | Intelligibility at the intended quiet delivery |
| Shout | Get behind the stone pillar! | Energy without clipping or harsh peaks |
| Short response | Understood. | A useful start and finish on a very short line |
| Long instruction | Take the east stair, cross the empty gallery, and leave the blue seal beside the locked door. | Pacing and a complete final clause |
| Punctuation | You found it? You found it. | Question and statement separation |
| Character continuity | Welcome back, traveler. I kept your room. | A recognizable character in a calmer mood |
| Interruption | First, check the rope; then lower the basket; finally, signal with the lantern. | Clean cancellation and subtitle behavior |
For listening, hide the candidate names and keep playback levels comparable so loudness does not decide the winner. Ask reviewers for a concrete reason, such as a mispronounced name, an unclear instruction, or a delivery that conflicts with the scene. Keep acting preference separate from the pass/fail checks for correct words, complete audio, and reliable interruption.
Keep model terms and voice-file permissions separate
The Pocket TTS model card lists CC BY 4.0 and an access gate, and explicitly prohibits unauthorized impersonation or cloning. The separate Kyutai voice catalog has mixed terms: voice donations and selected Voice-Zero assets are listed as CC0; Alba MacKenna and VCTK material as CC BY 4.0; Expresso and EARS as non-commercial. The catalog includes merchant and announcer performances, but it is not one uniformly licensed voice pack.[3][4]
For each candidate, save the exact model revision and applicable terms, the voice file’s source and license, required credits, and evidence of permission for the intended use. Do not clone a real performer without explicit lawful consent. If a voice’s permissions are unclear, keep it out of the release candidate until resolved. A permissive model label does not settle every right in a recording or a person’s likeness.
Make the final choice with a small evidence record
Use the worksheet to record the chosen scene, acceptance budgets, reviewers’ reasons, failure clips, and the exact versions tested. The evidence file also maps each source claim to a decision it can support and a question it leaves unanswered. This is a documentation-based workflow, not a ranking, listening test, or production integration. No game engine, speech model, or Goblin3D generation was run for it.
Evidence used
- NPC voice decision audit and ten-line audition worksheet
Original source-linked decision matrix and blank, reusable audition records. No audio was generated and all result fields are null.
Sources
- 1.Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning — Hugging Face. Accessed 2026-10-02.
- 2.Open TTS Leaderboard: About, Listen, and Streaming — Hugging Face. Accessed 2026-10-02.
- 3.Pocket TTS model card and access conditions — Kyutai. Accessed 2026-10-02.
- 4.Kyutai TTS voice catalog and per-folder licenses — Kyutai. Accessed 2026-10-02.
Keep the prototype’s systems testable
Continue with the browser-game guide to separate scene assets, rendering, and gameplay before adding more systems.
Read the browser-game guideRelated Goblin3D guides and tools
Sources, product facts, and original evidence were checked before publication.