Skip to content

bin/vh pace: bring the pace of the lines in one narration take closer, without synthesizing again - #72

Merged
ZLHad merged 1 commit into
mainfrom
claude/narration-pace
Oct 7, 2026
Merged

ZLHad merged 1 commit into
mainfrom
claude/narration-pace

Conversation

@ZLHad

@ZLHad ZLHad commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Why

#70 added a paragraph to playbook/04-audio.md ("一条 take 里快慢差得明显时,可以逐句拉近") describing how a film's one-request Gemini take was evened out line by line with a project script, without synthesizing again. This makes it a command. The maintainer asked that it stay gentle ("别太严格这个语速"), so it is optional, moves each line only part of the way, leaves lines that are already close as they are, and its numbers are a reference for the ear; --dry-run shows them before anything is written. The name and flags were agreed before building: bin/vh pace, reference = the median of the lines by default, word times carried through the edit by default with optional re-transcription.

What changed

bin/vh pace <project> [zh|en] (tools/audio/pace.py; numpy through uv, plus mlx-whisper only with --align whisper)

  • Per line: consecutive lines are split at the silence nearest their ASR boundary, then the cut points are refined from the waveform: step back from the aligned start while the signal is within 32 dB of the take's speech level, so a first syllable the ASR placed late is not clipped; the same forward at the end.
  • Pauses inside a line longer than --pause (0.26 s) are shortened to it with 6 ms fades, unless the script asked for them (<short pause>, <long pause>, <#s#>).
  • Pace = spoken units per voiced second: CJK characters with digits as they are read (75 = 七十五, 1953年 = 5), English syllables. Reference: the median of the lines, --ref id / --ref first-last, or --target N. Tempo (reference / pace) ^ --strength (0.8) within --clamp (0.92–1.18) through ffmpeg atempo (pitch kept); lines within --tolerance (5%) and lines under 4 units or 0.5 s keep their pace.
  • Pauses between lines stay the take's own, voice to voice, clamped into a range by how the line ends (comma or none 0.12–0.40 s, full stop 0.25–0.70 s, end of a block 0.45–1.20 s); --extra id=s adds or takes away after a line. The silence before the first line (--lead) and after the last is kept. Joined in numpy, never anullsrc + concat.
  • Output: voiceover.<lang>.wav and timeline.<lang>.json written back (plus the timeline.json / voiceover.wav copies when they belong to that language), originals kept as .raw, then bin/vh captions <p> <lang>. A second run starts again from the .raw take (settings never compound); a new bin/vh tts take replaces the .raw (told apart by a sha256 of the paced file in the timeline's pace block); --restore puts the original back, and refuses when the current take is not the one pace wrote. The timeline is replaced before the wav, so a run stopped between the two can only end in a refusal, never in the paced file being taken for a new original. Beat-grid timelines and dialogues are refused.
  • Word times: --align map (default) carries every word through the cuts and the tempo change. --align gemini (one call per ≤ 9 min) or --align whisper (local mlx-whisper, Apple Silicon) transcribe the new take and place the lines with tts.py's align_lines, refresh asr and print how far they are from the mapped times; the same flag transcribes a take that has no word times yet. A failed transcription keeps the paced take with mapped times and exits 1.

Docs: tools/audio/README.md (new section 逐句拉近语速, with measurements), playbook/04-audio.md (the paragraph names the command; step 4 now keeps the take's own gaps inside ranges instead of re-laying them, and the word times are carried over), CLAUDE.md command list (+ AGENTS.md), bin/vh help, CHANGELOG.md.

CI: pace_checks in tools/ci.sh: a synthetic take of tone-burst "syllables" with known onsets (lines at 4–7 a second, an inner pause to cap and one the script asks for, gaps too short and too long, an ASR start 0.12 s late, and two lines sharing one ASR boundary inside the first line's last syllable). It checks against the known syllables, not the tool's own numbers: --dry-run writes nothing; all 42 syllables kept and every mapped word within 60 ms of its syllable; spread of paces smaller; tempos inside the clamp; the inner pause capped and the asked-for one kept; gaps in range; .raw = original; captions redone; a second run byte-identical; --restore byte-identical to the original; a beat-grid timeline refused. Mutations it catches: no step-back at line starts (a word 0.12 s off), mapping that ignores the tempo (0.27 s off), no silence search at line boundaries (a syllable split in two).

How it was verified

  • The synthetic take above, by hand as well: mapped word starts vs real onsets median 2–3 ms, max 35–38 ms (atempo does not stretch a line perfectly evenly), all 42 syllables present.
  • A real 2-minute, 27-line one-request Gemini take, on a scratch copy (the project's own files were only read): 5.18–7.38 chars/s (sd 0.50) → 5.88–6.83 (sd 0.21), 13 lines' tempo untouched, 129.5 s → 121.4 s. Every mapped line start is within 0.034 s of the voice onset in the new audio. --align whisper (large-v3-turbo): line starts 0.19 s early (median), max 0.27 s, which is why mapping is the default. --align gemini on the previous build of the same take (identical but for the boundary fix, 121.41 vs 121.42 s): line starts within 0.03 s (median) / 0.13 s (max) of the mapped ones; on the final build Gemini's daily transcription quota was spent, which exercised the failure path (take written, mapped times kept, pace.align stays map, exit 1). bin/vh qa click warnings: 57 on the original take, 53 after.
  • The real take found one bug the first synthetic take did not have: the aligner had put one line's end and the next one's start at the same instant, 20 ms inside the first line's last syllable, so the real 0.54 s gap was treated as a pause inside the second line. Lines are now split at the silence nearest the boundary, and CI has that case.
  • State handling by hand: re-run byte-identical and .raw = original; --dry-run leaves every file's hash unchanged; a new take replaces the .raw (said on stdout); a swapped wav under a paced timeline is refused; --restore refuses then, and works on the paced take (byte-identical, .raw removed); --ref unknown id / --extra unknown id / a bad --clamp are errors.
  • Built on Lessons from a narrated film with 3D: one voice in a one-request take, labels at the floor from the start, text swaps and readouts, glass in Three.js #70's branch, then replayed onto main after Lessons from a narrated film with 3D: one voice in a one-request take, labels at the floor from the start, text swaps and readouts, glass in Three.js #70 and bin/vh tts: a 429 that names a per-day Gemini quota fails at once #71 were merged (one CHANGELOG conflict: both entries kept). tools/ci.sh --committed passes on the result, also with VH_BASH=/bin/bash (shellcheck and pyflakes through the uv shim).

…, without synthesizing again

A one-request take can run audibly fast and slow from line to line. playbook/04 described how to even it out on the
take (#70); this makes it a command. It stays gentle, as the maintainer asked ("别太严格这个语速"): optional, each line
moves part of the way towards the median (or --ref / --target), lines already close keep their pace, and --dry-run
shows the numbers before anything is written.

Lines split at the silence nearest their ASR boundary; cut points refined from the waveform; long pauses inside a line
capped unless the script asked for them; ffmpeg atempo per line; the take's own pauses between lines kept inside a
range by punctuation and block. Word times carried through the edit by default (or transcribed again with gemini or a
local whisper, through tts.py's align_lines). The originals stay as .raw; a re-run starts from them, --restore puts
them back. tools/ci.sh checks it on a synthetic take with known syllable onsets.
@ZLHad
ZLHad marked this pull request as ready for review October 7, 2026 15:31
@ZLHad
ZLHad enabled auto-merge (squash) October 7, 2026 15:32
@ZLHad
ZLHad merged commit 5d7d32e into main Oct 7, 2026
2 checks passed
@ZLHad
ZLHad deleted the claude/narration-pace branch October 7, 2026 15:35
ZLHad added a commit that referenced this pull request Oct 7, 2026
* bin/vh pace: fixes from the review of #72

An independent review of #72 (merged before its findings were in) found that an ASR boundary off by more than about
0.1 s could drop syllables or move one behind a pause, a crash when a line has no audio, a missing GEMINI_API_KEY found
only after the take was written, and smaller gaps (--restore --dry-run restored; --ref on a "……" line divided by zero;
a 22.05 kHz hop drift; "……" lines collapsed; the .raw timeline rewritten with a transcription).

A line's voice is now everything voiced between the silences that part it from its neighbours; the ASR start and end
only locate those silences, so nothing voiced can be dropped. Word times follow the audio through atempo's uneven
stretch (loudness envelopes aligned by DTW). Keys and mlx-whisper are checked before anything is written. tools/ci.sh
covers each case; run against #72's pace.py the new checks fail where the review said.

* bin/vh pace: fixes from the second review (of #73)

An independent review of this PR found:
- the breath rule peeled every short isolated syllable at a line edge with no loudness test (a line opening with three
  short loud syllables measured 6.36 for a true 5.07); now one blip per side, 15 dB under the line's speech;
- a line with no audio in the middle got zero-width times and an extra clamped pause, and a whispered line was replaced
  by silence; now such a line stays inside the pause around it, that stretch of the take is copied as it is, the line
  gets a 0.1 s span and is named;
- spans that go back in time (a hand edit) repeated up to 3 s of audio; now refused;
- a partial script.txt made a block break the note denied; lines missing from it now count as one block;
- CI's 60 ms word bound could not catch a revert of the DTW alignment (now 35 ms: 27 with it, 61 without), and the
  breath rule, the cache, --target/--extra and the notes had no cases;
- the dense DTW matrix took 1.2 GB for one 120 s line (now banded, 0.23 GB); the first transcription was unguarded;
  an unreachable 50 ms guard.

Each new CI case fails against the previous commit and passes now.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant