Repository navigation
bin/vh pace: bring the pace of the lines in one narration take closer, without synthesizing again - #72
Merged
Conversation
…, without synthesizing again A one-request take can run audibly fast and slow from line to line. playbook/04 described how to even it out on the take (#70); this makes it a command. It stays gentle, as the maintainer asked ("别太严格这个语速"): optional, each line moves part of the way towards the median (or --ref / --target), lines already close keep their pace, and --dry-run shows the numbers before anything is written. Lines split at the silence nearest their ASR boundary; cut points refined from the waveform; long pauses inside a line capped unless the script asked for them; ffmpeg atempo per line; the take's own pauses between lines kept inside a range by punctuation and block. Word times carried through the edit by default (or transcribed again with gemini or a local whisper, through tts.py's align_lines). The originals stay as .raw; a re-run starts from them, --restore puts them back. tools/ci.sh checks it on a synthetic take with known syllable onsets.
ZLHad
marked this pull request as ready for review
October 7, 2026 15:31
ZLHad
enabled auto-merge (squash)
October 7, 2026 15:32
ZLHad
added a commit
that referenced
this pull request
Oct 7, 2026
* bin/vh pace: fixes from the review of #72 An independent review of #72 (merged before its findings were in) found that an ASR boundary off by more than about 0.1 s could drop syllables or move one behind a pause, a crash when a line has no audio, a missing GEMINI_API_KEY found only after the take was written, and smaller gaps (--restore --dry-run restored; --ref on a "……" line divided by zero; a 22.05 kHz hop drift; "……" lines collapsed; the .raw timeline rewritten with a transcription). A line's voice is now everything voiced between the silences that part it from its neighbours; the ASR start and end only locate those silences, so nothing voiced can be dropped. Word times follow the audio through atempo's uneven stretch (loudness envelopes aligned by DTW). Keys and mlx-whisper are checked before anything is written. tools/ci.sh covers each case; run against #72's pace.py the new checks fail where the review said. * bin/vh pace: fixes from the second review (of #73) An independent review of this PR found: - the breath rule peeled every short isolated syllable at a line edge with no loudness test (a line opening with three short loud syllables measured 6.36 for a true 5.07); now one blip per side, 15 dB under the line's speech; - a line with no audio in the middle got zero-width times and an extra clamped pause, and a whispered line was replaced by silence; now such a line stays inside the pause around it, that stretch of the take is copied as it is, the line gets a 0.1 s span and is named; - spans that go back in time (a hand edit) repeated up to 3 s of audio; now refused; - a partial script.txt made a block break the note denied; lines missing from it now count as one block; - CI's 60 ms word bound could not catch a revert of the DTW alignment (now 35 ms: 27 with it, 61 without), and the breath rule, the cache, --target/--extra and the notes had no cases; - the dense DTW matrix took 1.2 GB for one 120 s line (now banded, 0.23 GB); the first transcription was unguarded; an unreachable 50 ms guard. Each new CI case fails against the previous commit and passes now.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
#70 added a paragraph to
playbook/04-audio.md("一条 take 里快慢差得明显时,可以逐句拉近") describing how a film's one-request Gemini take was evened out line by line with a project script, without synthesizing again. This makes it a command. The maintainer asked that it stay gentle ("别太严格这个语速"), so it is optional, moves each line only part of the way, leaves lines that are already close as they are, and its numbers are a reference for the ear;--dry-runshows them before anything is written. The name and flags were agreed before building:bin/vh pace, reference = the median of the lines by default, word times carried through the edit by default with optional re-transcription.What changed
bin/vh pace <project> [zh|en](tools/audio/pace.py; numpy through uv, plus mlx-whisper only with--align whisper)--pause(0.26 s) are shortened to it with 6 ms fades, unless the script asked for them (<short pause>,<long pause>,<#s#>).--ref id/--ref first-last, or--target N. Tempo(reference / pace) ^ --strength(0.8) within--clamp(0.92–1.18) through ffmpegatempo(pitch kept); lines within--tolerance(5%) and lines under 4 units or 0.5 s keep their pace.--extra id=sadds or takes away after a line. The silence before the first line (--lead) and after the last is kept. Joined in numpy, neveranullsrc+concat.voiceover.<lang>.wavandtimeline.<lang>.jsonwritten back (plus thetimeline.json/voiceover.wavcopies when they belong to that language), originals kept as.raw, thenbin/vh captions <p> <lang>. A second run starts again from the.rawtake (settings never compound); a newbin/vh ttstake replaces the.raw(told apart by a sha256 of the paced file in the timeline'spaceblock);--restoreputs the original back, and refuses when the current take is not the onepacewrote. The timeline is replaced before the wav, so a run stopped between the two can only end in a refusal, never in the paced file being taken for a new original. Beat-grid timelines and dialogues are refused.--align map(default) carries every word through the cuts and the tempo change.--align gemini(one call per ≤ 9 min) or--align whisper(local mlx-whisper, Apple Silicon) transcribe the new take and place the lines withtts.py'salign_lines, refreshasrand print how far they are from the mapped times; the same flag transcribes a take that has no word times yet. A failed transcription keeps the paced take with mapped times and exits 1.Docs:
tools/audio/README.md(new section 逐句拉近语速, with measurements),playbook/04-audio.md(the paragraph names the command; step 4 now keeps the take's own gaps inside ranges instead of re-laying them, and the word times are carried over),CLAUDE.mdcommand list (+AGENTS.md),bin/vhhelp,CHANGELOG.md.CI:
pace_checksintools/ci.sh: a synthetic take of tone-burst "syllables" with known onsets (lines at 4–7 a second, an inner pause to cap and one the script asks for, gaps too short and too long, an ASR start 0.12 s late, and two lines sharing one ASR boundary inside the first line's last syllable). It checks against the known syllables, not the tool's own numbers:--dry-runwrites nothing; all 42 syllables kept and every mapped word within 60 ms of its syllable; spread of paces smaller; tempos inside the clamp; the inner pause capped and the asked-for one kept; gaps in range;.raw= original; captions redone; a second run byte-identical;--restorebyte-identical to the original; a beat-grid timeline refused. Mutations it catches: no step-back at line starts (a word 0.12 s off), mapping that ignores the tempo (0.27 s off), no silence search at line boundaries (a syllable split in two).How it was verified
--align whisper(large-v3-turbo): line starts 0.19 s early (median), max 0.27 s, which is why mapping is the default.--align geminion the previous build of the same take (identical but for the boundary fix, 121.41 vs 121.42 s): line starts within 0.03 s (median) / 0.13 s (max) of the mapped ones; on the final build Gemini's daily transcription quota was spent, which exercised the failure path (take written, mapped times kept,pace.alignstaysmap, exit 1).bin/vh qaclick warnings: 57 on the original take, 53 after..raw= original;--dry-runleaves every file's hash unchanged; a new take replaces the.raw(said on stdout); a swapped wav under a paced timeline is refused;--restorerefuses then, and works on the paced take (byte-identical,.rawremoved);--refunknown id /--extraunknown id / a bad--clampare errors.mainafter Lessons from a narrated film with 3D: one voice in a one-request take, labels at the floor from the start, text swaps and readouts, glass in Three.js #70 and bin/vh tts: a 429 that names a per-day Gemini quota fails at once #71 were merged (one CHANGELOG conflict: both entries kept).tools/ci.sh --committedpasses on the result, also withVH_BASH=/bin/bash(shellcheck and pyflakes through the uv shim).