Whisper Word Timestamps to Reviewable Game Subtitles

Turn estimated Whisper word timings into reviewable game-dialogue cues with a reproducible Python converter, WebVTT export, and explicit timing checks.

colinkoko6 min read

Key takeaways

  • Treat estimated word timestamps as review input; keep questionable timing visible.
  • A 40-character, two-line starting policy still needs manual line-break and in-game layout review.
  • Retain clip-relative milliseconds and source word indices, then supply speaker and meaningful sound information separately.

Start with estimated word timings, then review the dialogue

To make game subtitles from Whisper output, request word timestamps, validate their order and duration, group the words into readable cues, and review those cues against the recording and your game UI. Keep the raw timing data alongside the edited subtitle asset so corrections remain traceable.[3]

A fantasy archivist beside a luminous waveform and three panels representing grouped subtitle cues.
Concept illustration of speech becoming subtitle cues. The waveform and panels are illustrative, not a software screenshot or measured timing result.

Whisper is a publicly released speech-recognition project. Its large-v3-turbo model card on Hugging Face documents local-file transcription and the return_timestamps="word" option. This guide focuses on the English transcription handoff, after a transcript exists. It makes no recommendation about the best model for your recording.[1][2]

Hugging Face describes Whisper word timestamps as an approximation produced using dynamic time warping over cross-attention weights. The pipeline returns text chunks with start/end pairs in seconds. Those estimates need listening review before they become an accepted dialogue timeline.[3]

Validate the input before grouping or rounding it

A missing end time should stop the handoff. Filling it with the next start would make a guessed alignment look authoritative. The supplied adapter rejects missing, nonnumeric, nonfinite, negative, reversed, and overlapping intervals under its explicitly single-speaker contract. It preserves input order instead of sorting mistakes away.

Millisecond conversion needs its own check. Two distinct second values can round to the same millisecond, leaving a zero-duration word. Validate both the original values and the quantized interval. Keep any failure attached to the input index so an editor can return to the exact word.

WebVTT requires each cue to end after it starts and lists cues in start-time order. It permits overlapping subtitle cues. Rejecting overlap here is a deliberately narrower adapter policy; simultaneous speakers need a different editorial and display design. The cited WebVTT document is a Candidate Recommendation Draft.[4]

Group words with a visible, adjustable policy

Microsoft’s game guidance recommends avoiding lines longer than 40 characters, generally showing at most two subtitle lines, and choosing editorially sensible line breaks manually when possible. Those recommendations motivate this adapter’s starting limits. A character count cannot establish the width of a line in your font.[5]

Use sentence punctuation, a configured pause, and line capacity to propose cue boundaries. Preserve the first included word’s start and the last included word’s end. When a cue is too brief or text-heavy, flag it for review rather than extending it into another line of speech. Timing and reading-speed thresholds in this download are configurable examples, not accessibility standards.

  • An overlong fantasy name triggers an explicit rejection for manual review. The adapter never truncates it or silently divides it across lines.
  • Check punctuation around abbreviations and quotations. A simple sentence-ending rule cannot understand every dramatic pause.
  • The adapter trims outside ASCII spaces and joins accepted word chunks with one space. Keep punctuation attached to its word and review this normalization; it is not a lossless transcript or multilingual segmenter.
  • Escape literal angle brackets and ampersands when serializing WebVTT so dialogue cannot accidentally become cue markup.[4]

Reproduce the synthetic subtitle handoff

Download and extract the evidence package below. It contains the original adapter, input fixtures, executable tests, a generated WebVTT file, and a JSON cue manifest. Run it locally with Python; the adapter makes no network requests and requires no model download.

python3 subtitle_adapter.py fixtures/dialogue.json --vtt review.vtt --manifest review.json
python3 test_adapter.py
Run after extracting the package. The first command writes new local review files; the second runs adapter tests, not an ASR model.

The recorded run on Python 3.12.14 passed 50 tests. Across three synthetic fixtures, 65 word chunks produced 15 cues, with exact source-index coverage and four cues flagged for review. Five separate invalid-input demonstrations were rejected. These are adapter results on authored data.

FixtureWords → cuesReview result
Dialogue41 → 7No flagged cues; longest line 39 code points
Review edges21 → 54 flagged cues; up to 2 lines, longest line 40
Exact milliseconds3 → 3Half-up rounding and an hour-scale clock preserved
Observed outputs from the bundled synthetic fixtures. Review flags use this adapter’s example thresholds.

Inspect the raw input and the exported cues side by side. The useful check is whether every accepted input word appears exactly once, whether each cue retains its real fixture boundaries, and whether the output keeps review flags visible. The saved output can be compared with a fresh run; a successful conversion is only the start of dialogue review.

Carry clip identity and timing into the game

Keep time relative to the audio clip that produced the transcript. A cutscene offset belongs in an explicit playback mapping; do not quietly add it to only one export. WebVTT timestamps describe offsets on the associated media timeline. A separate manifest lets a game importer retain this relationship without depending on subtitle text as an identifier.[4]

FieldPurpose
Clip identityWhich recording and revision do these timings belong to?
Start/end millisecondsWhen does the cue apply on that recording?
Source word indicesWhich original chunks formed this cue?
Cue identifierWhich cue in this particular export is being reviewed?
Speaker and contextWho speaks, and what information must the player see?
Review flagsWhich timing or reading issues still need a decision?
Suggested handoff fields and the question each answers.

Ordinal cue identifiers are useful inside one export, but they can change when words are regrouped. Keep your production dialogue key outside that sequence. Associate the final cue with a recording revision, script line, speaker supplied by the authoring system, and localization entry in your own importer. This article does not supply or test an engine-specific integration.

Finish the review in real game context

Listen against the recording and script before accepting proper names, repeated words, and speech during pauses. The model card warns that Whisper can output unspoken or repetitive text. Timestamp validation cannot detect those content errors.[2]

Speech subtitles cover dialogue. Microsoft distinguishes them from captions that also communicate meaningful non-speech sound. Plan separate authored information for an offscreen warning, a door knock, music that changes the scene’s meaning, or speaker context. A transcript converter does not supply that information.[5]

In your game, review the smallest supported screen, an enlarged subtitle setting, busy backgrounds, dialogue interruptions, and rapid scene changes. Confirm that a subtitle is cleared or rescheduled when its associated audio is cancelled. These are proposed acceptance checks, not completed tests of the downloadable fixture.

Know the limits before adopting the converter

The adapter is a small, inspectable starting point for English single-speaker material. It does not transcribe audio, infer speakers, translate dialogue, align phonemes, create lip-sync data, or certify accessibility. Its layout limits count code points, not shaped glyph widths. Complex writing systems, overlapping actors, manual retiming, and renderer-specific cue placement need additional work.

Start with one short recording you have permission to use. Keep the transcript, review corrections, cue manifest, and final audio revision together. Once the complete handoff works for that clip, test more varied dialogue before applying it to a larger scene library.

Evidence used

Sources

  1. 1.Introducing Whisper: public speech-recognition project — OpenAI. Accessed 2026-10-04.
  2. 2.Whisper large-v3-turbo model card and word-timestamp example — OpenAI. Accessed 2026-10-04.
  3. 3.Transformers v5.17.0 ASR pipeline timestamp contract — Hugging Face. Accessed 2026-10-04.
  4. 4.WebVTT Candidate Recommendation Draft, May 20, 2026 — W3C. Accessed 2026-10-04.
  5. 5.Xbox Accessibility Guideline 104: subtitles and captions — Microsoft. Accessed 2026-10-04.

Plan the visual side of your dialogue scene

Explore Goblin3D’s asset workflows while you prepare the character and scene references for your prototype.

Explore Goblin3D features

Sources, product facts, and original evidence were checked before publication.

Share your feedback

Sign in to share feedback with the Goblin3D team.

Sign in

support@goblin3d.ai