# Word-timestamp subtitle cue evidence pack

An original, runnable Python standard-library adapter for **single-speaker English dialogue represented as space-separated word chunks**. It converts authored word/timestamp fixtures to WebVTT and a game-neutral JSON manifest. It makes no network requests, loads no model, needs no token, and calls no game engine.

## Evidence boundary

Every fixture in this pack is synthetic. The dialogue, times, code, and tests were authored for the example. No audio was transcribed. The test results establish behavior of this particular adapter on these inputs, not Whisper accuracy, alignment accuracy, inference performance, production reliability, engine compatibility, or player accessibility. A passing test is not a claim that the subtitles are ready to ship.

The result is an editorial starting point. Check the transcript against the actual recording, word timing against the waveform, punctuation and phrase breaks, reading comfort, and the target game's font, UI scale, viewport, background contrast, safe area, and localization. Resolve intentional overlapping speech with a suitable multi-speaker workflow rather than deleting or shifting words to satisfy this adapter.

## Run locally

Python 3.10 or later is required. No installation step or third-party package is needed. From this directory:

```sh
python3 subtitle_adapter.py fixtures/dialogue.json \
  --vtt results/dialogue.vtt \
  --manifest results/dialogue.manifest.json
python3 test_adapter.py
python3 run_evidence.py
```

`run_evidence.py` regenerates the three fixture outputs, rejection demonstrations, test log, summary, checksums, and `subtitle-cue-evidence.zip`. It does not use the host's local clock. For a new recorded run, optionally pass `--checked-at-utc` with an independently verified ISO-8601 UTC timestamp. Without it, the summary records no evidence timestamp. Do not reuse the published timestamp for a later run. Test elapsed time in the log can vary; it is not an ASR benchmark. Converter VTT/JSON outputs are byte-deterministic for identical input and configuration. ZIP contents include run-specific evidence, so the full archive is not promised byte-identical across test runs.

CLI outputs overwrite the specified output files; use new paths if you want to keep previous results. Parent directories must exist. Input and output paths must be distinct, including existing hard links. Input validation finishes before writing outputs. An operating-system write failure can still leave just one of the two outputs; writing the pair is not a filesystem transaction. Invalid input exits with status 2 and an error message.

## Strict input contract

Supply a JSON array of chunks, or an object whose `chunks` member is that array. Other envelope/chunk keys are ignored. The optional envelope `text` is not used. Each chunk must have one textual token and exactly two numerical timestamps in **seconds**:

```json
{"chunks": [
  {"text": " Hold", "timestamp": [0.125, 0.6]},
  {"text": " steady.", "timestamp": [0.6, 1.25]}
]}
```

- Times are offsets from the start of the associated clip. Zero means that clip's start. No global timeline offset, audio trimming correction, clip lookup, or automatic speaker inference is performed. Clip identity and speaker metadata belong to your integration; this manifest does not provide them.
- The adapter trims ASCII spaces at each token's edges and joins normalized tokens with one ASCII space. The unmodified `source_text` remains in the manifest for audit. This is **not** a lossless raw-ASR transcript adapter. It does not preserve original inter-word spacing.
- A chunk containing internal whitespace is rejected. Segment-level timestamps cannot be distributed across multiple words without additional alignment evidence. Tabs, newlines, nonbreaking spaces, control characters, and lone surrogates are rejected.
- Punctuation is expected to be attached to a word. A punctuation-only token is accepted as a token but is not merged with a neighboring word; such input needs upstream editorial normalization. Abbreviations such as `Dr.` can create undesirable breaks. There is no language-aware sentence parser.
- Missing/null/string/bool/non-finite timestamps, negative times, end less than or equal to start, unsorted words, and overlapping words are rejected. Adjacent words may touch (`next.start == previous.end`). No time is fabricated to fill missing data.
- Overlap rejection is this **single-speaker adapter policy**, not a WebVTT prohibition. WebVTT permits overlapping cues.
- Empty input is allowed and emits an empty WebVTT document and manifest.
- Bounds: UTF-8 input file up to 2 MiB; up to 10,000 chunks; each time from 0 through 86,400 seconds. Duplicate JSON keys and nonstandard `NaN`/`Infinity` constants are rejected. The file-size bound applies to the CLI reader; direct Python callers supply already-loaded objects.

## Milliseconds and timing preservation

JSON fractional literals are parsed with `Decimal`. Convert seconds to integer milliseconds by decimal round-half-up: `0.0005` seconds becomes 1 ms, while `1.234` seconds stays 1,234 ms. This is not Python's default ties-to-even `round`. When calling the Python API directly with floats, conversion uses the float's `str` representation; use JSON decimals or `Decimal` for a precisely chosen decimal value.

Validate source ordering and non-overlap **before** rounding, so a tiny source overlap cannot disappear unnoticed. Validate each rounded end remains greater than its rounded start. For example, `[0.0001, 0.0004]` collapses to `[0, 0]` and is rejected; the adapter does not add a millisecond to make it pass. The manifest retains original numeric seconds as strings and the rounded integer milliseconds.

Each cue begins at its first included word's rounded start and ends at its last included word's rounded end. There is no lead-in, hold-time extension, duration padding, truncation, or removal of words. Silence inside a grouped cue remains inside that cue's duration. Clock formatting uses integer division and supports hours, including timestamps beyond one hour. The converter's input-time bound is separate from the generic clock formatter.

## Grouping, layout, and review flags

The greedy algorithm makes a new cue when the preceding token ends in `. ! ? , ; :` (including a limited set of closing quotes/brackets), the next word follows a gap of at least 600 rounded milliseconds, or whole-word wrapping would exceed line capacity. A final unfinished group closes at end of input. The manifest records which rule closed each cue.

Defaults are 40 Unicode codepoints per line and at most 2 lines. A single word longer than the configured limit is **rejected for manual review**, never truncated or divided into invented timed fragments. Word wrapping counts normalized plain text before VTT escaping, not the longer entity strings.

Python's `len` counts Unicode codepoints, not grapheme clusters, displayed glyphs, reading units, or pixels. For example, an `e` plus combining acute accent counts as 2, and an emoji with a joiner can count as several. The fixtures and grouping target space-separated English; this is not a validated multilingual layout or shaping solution. No font measurement is performed. Inspect the actual rendered text.

Configurable example thresholds flag a cue if its duration is below 1,000 ms, above 6,000 ms, or its character rate exceeds 20 codepoints/second. **These are illustrative review triggers, not universal standards, XAG requirements, or pass/fail accessibility criteria.** Equality at the thresholds is not flagged. A rate is calculated from the normalized cue text, including inter-word spaces and punctuation, divided by the complete cue duration. It uses full precision for comparisons and is displayed rounded to 3 decimals. Thresholds do not change the timing or force extra cue splits.

All settings are exposed as CLI options:

```sh
python3 subtitle_adapter.py fixtures/dialogue.json \
  --vtt results/custom.vtt --manifest results/custom.manifest.json \
  --line-codepoints 36 --max-lines 2 --pause-ms 700 \
  --min-duration-ms 900 --max-duration-ms 6500 --max-cps 18
```

Accepted configuration ranges: line codepoints 1–120; lines 1–2; pause 1–60,000 ms; each duration 1–86,400,000 ms with minimum no greater than maximum; CPS 1–1,000. These ranges are implementation bounds, not recommended display settings.

## Output and integration

`cue-0001`, `cue-0002`, and so on are deterministic **ordinal IDs** for the same input/configuration. They are not durable content IDs: editing or resegmentation can change which text has a given ordinal. Store a clip identity and a separate durable editorial key in your integration when needed.

Each cue includes integer start/end/duration, normalized unwrapped text, wrapped lines, zero-based source word indices, boundary reason, codepoint count, rate, and review flags. The manifest's `words` array preserves each token's source text, original numerical seconds, and rounded times. Source indices let reviewers audit coverage within the same input; they are not stable across edits to that input.

WebVTT payloads escape `&`, `<`, and `>` as literal text, including text that resembles tags or a timing arrow. The JSON retains plain text; a consuming engine must apply its own rich-text/markup escaping. The pack does not validate an engine importer or introduce speaker styling, captions for meaningful non-speech audio, translation, or persistence storage.

## Files and tests

- `subtitle_adapter.py`: validated converter and CLI
- `test_adapter.py`: unit/integration tests, including 100 seeded synthetic schedules inside one invariant test
- `run_evidence.py`: local evidence/ZIP builder
- `fixtures/dialogue.json`: authored single-speaker dialogue
- `fixtures/review_edges.json`: capacity, fast display, long display, and literal-markup cases
- `fixtures/exact_milliseconds.json`: half-up rounding, exact milliseconds, and >1-hour clocks
- `results/*.vtt` and `results/*.manifest.json`: deterministic expected fixture outputs
- `results/rejection-cases.json`: five rejected-input demonstrations
- `results/unittest.txt`, `results/summary.json`, `results/checksums.json`: actual test evidence and SHA-256 hashes
- `source-provenance.json`: sources and the limited claims they support

Tests cover malformed structures, missing values, timing ordering, source overlap before quantization, millisecond collapse, half-up ties, high-precision decimal boundaries, Unicode codepoint behavior, escaping, full word-index coverage, no duplicate/dropped tokens, exact expected VTT bytes, CLI success/failure, deterministic serialization, and integer hour formatting. These are targeted tests, not formal verification or independent WebVTT conformance certification.

## Primary references

The following sources informed the interface and constraints; no third-party article, code, audio, model, or artwork is redistributed in this pack.

- [Hugging Face ASR pipeline documentation](https://huggingface.co/docs/transformers/en/main_classes/pipelines#transformers.AutomaticSpeechRecognitionPipeline), also checked at [v5.17.0](https://huggingface.co/docs/transformers/v5.17.0/en/main_classes/pipelines#transformers.AutomaticSpeechRecognitionPipeline): documents word timestamp chunks and describes Whisper's DTW-based word timestamp approximation. This pack accepts a deliberately narrower input subset; it does not call that API.
- [W3C WebVTT](https://www.w3.org/TR/webvtt1/), also checked at the [20 May 2026 Candidate Recommendation Draft](https://www.w3.org/TR/2026/CRD-webvtt1-20260520/): cue ends must follow starts, and distinct cues may overlap. This adapter's non-overlap input policy is stricter.
- [Microsoft Xbox Accessibility Guideline 104](https://learn.microsoft.com/en-us/xbox/accessibility/xbox-accessibility-guidelines/104): advises avoiding lines over 40 characters and generally using at most two subtitle lines, with exceptional cases for three. The adapter's codepoint counter is a pragmatic proxy, not a proof of compliance with the guideline's broader visual requirements.
