Key takeaways
- Treat estimated word timestamps as review input; keep questionable timing visible.
- A 40-character, two-line starting policy still needs manual line-break and in-game layout review.
- Retain clip-relative milliseconds and source word indices, then supply speaker and meaningful sound information separately.
Start with estimated word timings, then review the dialogue
To make game subtitles from Whisper output, request word timestamps, validate their order and duration, group the words into readable cues, and review those cues against the recording and your game UI. Keep the raw timing data alongside the edited subtitle asset so corrections remain traceable.[3]

Whisper is a publicly released speech-recognition project. Its large-v3-turbo model card on Hugging Face documents local-file transcription and the return_timestamps="word" option. This guide focuses on the English transcription handoff, after a transcript exists. It makes no recommendation about the best model for your recording.[1][2]
Hugging Face describes Whisper word timestamps as an approximation produced using dynamic time warping over cross-attention weights. The pipeline returns text chunks with start/end pairs in seconds. Those estimates need listening review before they become an accepted dialogue timeline.[3]
Validate the input before grouping or rounding it
A missing end time should stop the handoff. Filling it with the next start would make a guessed alignment look authoritative. The supplied adapter rejects missing, nonnumeric, nonfinite, negative, reversed, and overlapping intervals under its explicitly single-speaker contract. It preserves input order instead of sorting mistakes away.
Millisecond conversion needs its own check. Two distinct second values can round to the same millisecond, leaving a zero-duration word. Validate both the original values and the quantized interval. Keep any failure attached to the input index so an editor can return to the exact word.
WebVTT requires each cue to end after it starts and lists cues in start-time order. It permits overlapping subtitle cues. Rejecting overlap here is a deliberately narrower adapter policy; simultaneous speakers need a different editorial and display design. The cited WebVTT document is a Candidate Recommendation Draft.[4]
Group words with a visible, adjustable policy
Microsoft’s game guidance recommends avoiding lines longer than 40 characters, generally showing at most two subtitle lines, and choosing editorially sensible line breaks manually when possible. Those recommendations motivate this adapter’s starting limits. A character count cannot establish the width of a line in your font.[5]
Use sentence punctuation, a configured pause, and line capacity to propose cue boundaries. Preserve the first included word’s start and the last included word’s end. When a cue is too brief or text-heavy, flag it for review rather than extending it into another line of speech. Timing and reading-speed thresholds in this download are configurable examples, not accessibility standards.
- An overlong fantasy name triggers an explicit rejection for manual review. The adapter never truncates it or silently divides it across lines.
- Check punctuation around abbreviations and quotations. A simple sentence-ending rule cannot understand every dramatic pause.
- The adapter trims outside ASCII spaces and joins accepted word chunks with one space. Keep punctuation attached to its word and review this normalization; it is not a lossless transcript or multilingual segmenter.
- Escape literal angle brackets and ampersands when serializing WebVTT so dialogue cannot accidentally become cue markup.[4]
Reproduce the synthetic subtitle handoff
Download and extract the evidence package below. It contains the original adapter, input fixtures, executable tests, a generated WebVTT file, and a JSON cue manifest. Run it locally with Python; the adapter makes no network requests and requires no model download.
python3 subtitle_adapter.py fixtures/dialogue.json --vtt review.vtt --manifest review.json
python3 test_adapter.pyThe recorded run on Python 3.12.14 passed 50 tests. Across three synthetic fixtures, 65 word chunks produced 15 cues, with exact source-index coverage and four cues flagged for review. Five separate invalid-input demonstrations were rejected. These are adapter results on authored data.
| Fixture | Words → cues | Review result |
|---|---|---|
| Dialogue | 41 → 7 | No flagged cues; longest line 39 code points |
| Review edges | 21 → 5 | 4 flagged cues; up to 2 lines, longest line 40 |
| Exact milliseconds | 3 → 3 | Half-up rounding and an hour-scale clock preserved |
Inspect the raw input and the exported cues side by side. The useful check is whether every accepted input word appears exactly once, whether each cue retains its real fixture boundaries, and whether the output keeps review flags visible. The saved output can be compared with a fresh run; a successful conversion is only the start of dialogue review.
Carry clip identity and timing into the game
Keep time relative to the audio clip that produced the transcript. A cutscene offset belongs in an explicit playback mapping; do not quietly add it to only one export. WebVTT timestamps describe offsets on the associated media timeline. A separate manifest lets a game importer retain this relationship without depending on subtitle text as an identifier.[4]
| Field | Purpose |
|---|---|
| Clip identity | Which recording and revision do these timings belong to? |
| Start/end milliseconds | When does the cue apply on that recording? |
| Source word indices | Which original chunks formed this cue? |
| Cue identifier | Which cue in this particular export is being reviewed? |
| Speaker and context | Who speaks, and what information must the player see? |
| Review flags | Which timing or reading issues still need a decision? |
Ordinal cue identifiers are useful inside one export, but they can change when words are regrouped. Keep your production dialogue key outside that sequence. Associate the final cue with a recording revision, script line, speaker supplied by the authoring system, and localization entry in your own importer. This article does not supply or test an engine-specific integration.
Finish the review in real game context
Listen against the recording and script before accepting proper names, repeated words, and speech during pauses. The model card warns that Whisper can output unspoken or repetitive text. Timestamp validation cannot detect those content errors.[2]
Speech subtitles cover dialogue. Microsoft distinguishes them from captions that also communicate meaningful non-speech sound. Plan separate authored information for an offscreen warning, a door knock, music that changes the scene’s meaning, or speaker context. A transcript converter does not supply that information.[5]
In your game, review the smallest supported screen, an enlarged subtitle setting, busy backgrounds, dialogue interruptions, and rapid scene changes. Confirm that a subtitle is cleared or rescheduled when its associated audio is cancelled. These are proposed acceptance checks, not completed tests of the downloadable fixture.
Know the limits before adopting the converter
The adapter is a small, inspectable starting point for English single-speaker material. It does not transcribe audio, infer speakers, translate dialogue, align phonemes, create lip-sync data, or certify accessibility. Its layout limits count code points, not shaped glyph widths. Complex writing systems, overlapping actors, manual retiming, and renderer-specific cue placement need additional work.
Start with one short recording you have permission to use. Keep the transcript, review corrections, cue manifest, and final audio revision together. Once the complete handoff works for that clip, test more varied dialogue before applying it to a larger scene library.
Evidence used
- Download the complete subtitle evidence pack
Original converter, tests, authored timing fixtures, generated WebVTT/JSON, source provenance, and actual test results. No audio or model weights.
- Read the adapter contract and reproduction instructions
Explicit input normalization, timing rules, scope limits, review policies, and commands.
- Inspect the original Python adapter
Dependency-free local word-chunk converter; does not transcribe audio or make network requests.
- Inspect the synthetic test summary
Actual fixture cue counts, review flags, and 50 passing adapter tests on invented data.
- Audit the dialogue cue manifest
Original source chunks and the seven proposed cues, preserving word indices and quantized timing.
- Inspect the executable adapter tests
Boundary, validation, escaping, deterministic output, CLI, and source-index-coverage tests.
Sources
- 1.Introducing Whisper: public speech-recognition project — OpenAI. Accessed 2026-10-04.
- 2.Whisper large-v3-turbo model card and word-timestamp example — OpenAI. Accessed 2026-10-04.
- 3.Transformers v5.17.0 ASR pipeline timestamp contract — Hugging Face. Accessed 2026-10-04.
- 4.WebVTT Candidate Recommendation Draft, May 20, 2026 — W3C. Accessed 2026-10-04.
- 5.Xbox Accessibility Guideline 104: subtitles and captions — Microsoft. Accessed 2026-10-04.
Plan the visual side of your dialogue scene
Explore Goblin3D’s asset workflows while you prepare the character and scene references for your prototype.
Explore Goblin3D featuresRelated Goblin3D guides and tools
Sources, product facts, and original evidence were checked before publication.