Repository navigation
Lessons from a narrated film with 3D: one voice in a one-request take, labels at the floor from the start, text swaps and readouts, glass in Three.js - #70
Merged
Conversation
…, labels at the floor from the start, text swaps and readouts, glass in Three.js The maintainer asked that the general lessons of a 2-minute narrated canvas + Three.js film go into the docs as short guidance. - playbook/04: in a one-request gemini take every line still carries its own style, so a per-line [direction] makes that line sound different; for one voice, no per-line directions. Evening out the pace line by line on one take (waveform-refined cuts, capped inner pauses, partial per-line atempo towards one target), described, not a command. When transcription is unavailable: the take is on disk; --resume later or align locally with whisper, mapping digits back to spelled-out years. - tools/audio/README: a daily-quota 429 was once waited on and retried like the per-minute limit, which looked hung. - playbook/03 §4: write labels at the floor from the first draft; a shot's headline is a 主标题; credits go into the description. - playbook/02 layer 3, TASTE_CHECKLIST #19: strips at text swaps and rolling readouts; a readout rests only on sourced values. - engines/README: RoomEnvironment reflections and the default LatheGeometry seam on glass; a fresnel rim shell instead.
- playbook/04: whisper output goes through tts.py's align_lines (it spreads spelled-out years over a heard "1953"; a home-made character aligner can misplace the line start); the one-voice advice is for single-narrator films and asks for one short overall style; the pace numbers are re-measured on the take with the same method (4.7–6.7 → 5.6–6.1 chars/s, sd 0.45 → 0.11); the evening steps say how the factor is computed, avoid anullsrc + concat, and write the take, the timeline and the captions back. - tools/audio/README: the quota text describes what the code does (a 429 under 90 s, or without a delay, is retried up to 5 times) and points to the playbook; the joined-synthesis section points to the one-voice advice. - playbook/03 §4: credits stay on screen where a licence or the type doc wants them (a data story's source line), as labels at the floor; the headline is "the line used as the shot's title". - TASTE_CHECKLIST #6: a shot's headline counts as a 主标题. - engines/README: "a row of glass objects"; dropping scene.environment also drops it for the other materials in that scene; the list's lead-in says most pitfalls come from the intro film. - playbook/02: "旧句退完、新句才进", as in #17. - CHANGELOG: rewritten to match, with the two floor changes said explicitly.
…umber The maintainer: "别太严格这个语速". The evening is optional (nothing to do when it sounds fine), aims for closer rather than equal, and the film's numbers are a reference.
ZLHad
marked this pull request as ready for review
October 7, 2026 14:30
ZLHad
added a commit
that referenced
this pull request
Oct 7, 2026
The ci.sh stub check sets no_proxy for 127.0.0.1, so an exported http_proxy cannot turn it red, and covers the fallback too: a 429 naming no daily quota that asks for 120 s fails at once. tools/audio/README says again that a 429 with no delay counts as 60 s; the #70 changelog entry points at the new one.
ZLHad
added a commit
that referenced
this pull request
Oct 7, 2026
* bin/vh tts: a 429 that names a per-day Gemini quota fails at once Gemini's 429 for a spent daily quota can ask for a retryDelay under a minute, and gemini_call took a 429 for the daily quota only when it asked for more than 90 s. Below that it waited and retried up to 5 times a call, so --align gemini looked hung once the daily quota was spent. gemini_call now reads the quotas the 429 body names: a google.rpc.QuotaFailure violation whose quotaId or quotaMetric has "PerDay" in it fails the call at once with a message naming the quota and the --resume recovery. Per-minute 429s are waited out and retried as before; a 429 naming no daily quota still fails at once only when it asks for more than 90 s. tools/ci.sh checks both cases against a stub server on 127.0.0.1. The quota sentences in tools/audio/README and playbook/04 say the daily quota now fails at once. * Fixes from the independent review The ci.sh stub check sets no_proxy for 127.0.0.1, so an exported http_proxy cannot turn it red, and covers the fallback too: a 429 naming no daily quota that asks for 120 s fails at once. tools/audio/README says again that a 429 with no delay counts as 60 s; the #70 changelog entry points at the new one.
ZLHad
added a commit
that referenced
this pull request
Oct 7, 2026
…, without synthesizing again (#72) A one-request take can run audibly fast and slow from line to line. playbook/04 described how to even it out on the take (#70); this makes it a command. It stays gentle, as the maintainer asked ("别太严格这个语速"): optional, each line moves part of the way towards the median (or --ref / --target), lines already close keep their pace, and --dry-run shows the numbers before anything is written. Lines split at the silence nearest their ASR boundary; cut points refined from the waveform; long pauses inside a line capped unless the script asked for them; ffmpeg atempo per line; the take's own pauses between lines kept inside a range by punctuation and block. Word times carried through the edit by default (or transcribed again with gemini or a local whisper, through tts.py's align_lines). The originals stay as .raw; a re-run starts from them, --restore puts them back. tools/ci.sh checks it on a synthetic take with known syllable onsets.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Lessons from a 2-minute narrated film (HyperFrames canvas + Three.js, one Gemini take for the whole narration), from its
LESSONS.md,DECISIONS.md, review notes and project scripts, written into the docs as short guidance without the film's own content.What goes where
playbook/04-audio.md§1 (旁白要导演): in a one-request gemini take (--join all) every line is still its own text item with its ownspeech_metadata.style(the overall style plus that line's[ ]), so a single line with a direction, often the opening, tends to sound different from the rest. For one voice in a single-narrator film: one short overall style, no per-line[ ].tools/audio/README.md"整段合成" points there.playbook/04-audio.md§1, after the pace paragraph: when lines in one take run audibly fast and slow, an optional way to bring them closer line by line without synthesizing again. Lines are meant to vary; nothing is to be done when it sounds fine, and the aim is closer, not one exact pace (the maintainer, mid-review: "别太严格这个语速"). Cuts come from the waveform (step back from the ASR start while the signal is loud: ASR line starts are approximate and often late when the first syllable is misheard, so a cut there clips it; the same forward at the end); pauses inside a line capped at about 0.25 s; each line'satempofactor is, for example, (reference ÷ its pace)^0.8, clamped to 0.92–1.18, the reference being a stretch the human accepted (e.g. the opening); gaps laid out again withoutanullsrc+concat; the take andtimeline.<lang>.jsonwritten back, aligned again,bin/vh captionsre-run. Described, not a command (see below).playbook/04-audio.md(词级时间和对稿检查): when transcription is unavailable (daily quota, persistent API errors).--align geminimay keep waiting and retrying on 429; the take is already on disk (voiceover.<lang>.wavper line,vo/<lang>/_blockNN.wavwith--join), so stop it and--resumelater, or take word times from a local whisper (mlx-whisperwhisper-large-v3-turboon Apple Silicon, names and terms ininitial_prompt) and align them to the script withtools/audio/tts.py'salign_lines, the aligner--align geminiuses. Whisper writes years as digits: similarity drops, and a home-made character aligner can misplace the line start (the film's own aligner put one late and the cut clipped the first character), so map them back to the script's spelling (一九五三) first.tools/audio/README.md(转写的配额, and the 90 s sentence in 词级时间): the old text said a spent daily quota fails at once. By the code, a 429 asking for under 90 s (or naming no delay, taken as 60 s) is retried like the per-minute limit, up to 5 times a call; that is what the film saw after its daily quota, and it looked hung.playbook/03-motion-design.md§4: write readable labels at the floor of the film'sWatch onfrom the first draft (44 px fordesktop); the film's first draft had about 40 at 24–40 px. The line used as a shot's title counts as a 主标题. Source, credit and asset lines go into the video description unless a licence or the type doc wants them on screen (a data story's source line), and then they are labels at the floor.playbook/02-verification.mdlayer 3: a strip at each text swap in one place (swatch: render on the CPU and encode single-threaded so every run gives the same bytes #17's old-out-before-new-in; the reviewer found three overlaps the author's own check missed) and at each rolling readout, whose resting values are checked against NOTES.engines/README.md(HyperFrames + Three.js pitfalls):RoomEnvironmentreflections drew a bright vertical strip down every glass flask, and a full-turnLatheGeometryseam faces a +z camera by default (phiStart = Math.PIturns it away). The glass scene dropsscene.environment(the other materials there lose it too; light them instead) and gets a fresnel rim shell: the same geometry scaled about 1.004, additiveShaderMaterial, no depth write,pow(1 − |N·V|, 2.6). The list's lead-in now says "most" of its pitfalls come from the intro film (Lessons from four films: copy that reads as written by a person, render and audio pitfalls, contributing notes #66 had already added others).CHANGELOG.md: an Unreleased entry.Floors touched (CONTRIBUTING rule 7)
templates/TASTE_CHECKLIST.mdDirector mode and decision-first review pages #19 (【底线】 half): names a readout resting on an interpolated value, between sourced values, as a fabricated number. The film's reviewer had already judged it under Director mode and decision-first review pages #19; the wording makes it explicit.phone,desktop) or 150 (feed) instead of a label's 44 / 80. The maintainer asked for this (the film's reviewer had failed five such lines at 52–72 px).Not in this PR
bin/vh tts: each its own PR if wanted.tools/audio/tts.pydecides "daily quota" only from the delay a 429 asks for; the quota id in the error body (…PerDay…) would say so directly. Reported separately, not changed here.Verified
tools/audio/tts.py: in a joined gemini request every line is atextitem with aspeech_metadataannotation whosestyleis the overall style joined with that line's direction; the per-line path writes the take before the first transcription call, and--joinwrites_blockNN.wavbefore transcribing it; a 429 asking for 90 s or less (60 s when it names no delay) is retried up to 5 times a call;align_linesspreads script units over a heard unit they do not match one to one ("一九五三" over "1953").LatheGeometry: vertexx = r·sin φ,z = r·cos φ, so the φ = 0 seam faces +z.tools/ci.sh --committedpasses, also withVH_BASH=/bin/bash; shellcheck and pyflakes ran through theuvrecipe. No keys, private paths or film names in the diff.Independent review (one Sonnet reviewer, read-only, fresh context): the code facts held. Fixed from it: the source-line advice contradicted type 05's on-screen source line (now: on screen when a licence or the type doc wants it, as a label at the floor); the year/whisper mechanism (the harness aligner does not start late; a home-made one can) and a pointer to
align_lines; the one-voice advice limited to single-narrator films and to one short style; the pace numbers, the factor formula, the 8-bit padding pitfall and writing the timeline and captions back; the two floor changes stated here and in the CHANGELOG, with #6 saying the headline rule itself; the README's quota text describing what the code does; small wording.