Skip to content

Lessons from a narrated film with 3D: one voice in a one-request take, labels at the floor from the start, text swaps and readouts, glass in Three.js - #70

Merged
ZLHad merged 3 commits into
mainfrom
claude/lessons-one-take-voice-labels
Oct 7, 2026
Merged

ZLHad merged 3 commits into
mainfrom
claude/lessons-one-take-voice-labels

Conversation

@ZLHad

@ZLHad ZLHad commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Lessons from a 2-minute narrated film (HyperFrames canvas + Three.js, one Gemini take for the whole narration), from its LESSONS.md, DECISIONS.md, review notes and project scripts, written into the docs as short guidance without the film's own content.

What goes where

  • playbook/04-audio.md §1 (旁白要导演): in a one-request gemini take (--join all) every line is still its own text item with its own speech_metadata.style (the overall style plus that line's [ ]), so a single line with a direction, often the opening, tends to sound different from the rest. For one voice in a single-narrator film: one short overall style, no per-line [ ]. tools/audio/README.md "整段合成" points there.
  • playbook/04-audio.md §1, after the pace paragraph: when lines in one take run audibly fast and slow, an optional way to bring them closer line by line without synthesizing again. Lines are meant to vary; nothing is to be done when it sounds fine, and the aim is closer, not one exact pace (the maintainer, mid-review: "别太严格这个语速"). Cuts come from the waveform (step back from the ASR start while the signal is loud: ASR line starts are approximate and often late when the first syllable is misheard, so a cut there clips it; the same forward at the end); pauses inside a line capped at about 0.25 s; each line's atempo factor is, for example, (reference ÷ its pace)^0.8, clamped to 0.92–1.18, the reference being a stretch the human accepted (e.g. the opening); gaps laid out again without anullsrc + concat; the take and timeline.<lang>.json written back, aligned again, bin/vh captions re-run. Described, not a command (see below).
  • playbook/04-audio.md (词级时间和对稿检查): when transcription is unavailable (daily quota, persistent API errors). --align gemini may keep waiting and retrying on 429; the take is already on disk (voiceover.<lang>.wav per line, vo/<lang>/_blockNN.wav with --join), so stop it and --resume later, or take word times from a local whisper (mlx-whisper whisper-large-v3-turbo on Apple Silicon, names and terms in initial_prompt) and align them to the script with tools/audio/tts.py's align_lines, the aligner --align gemini uses. Whisper writes years as digits: similarity drops, and a home-made character aligner can misplace the line start (the film's own aligner put one late and the cut clipped the first character), so map them back to the script's spelling (一九五三) first.
  • tools/audio/README.md (转写的配额, and the 90 s sentence in 词级时间): the old text said a spent daily quota fails at once. By the code, a 429 asking for under 90 s (or naming no delay, taken as 60 s) is retried like the per-minute limit, up to 5 times a call; that is what the film saw after its daily quota, and it looked hung.
  • playbook/03-motion-design.md §4: write readable labels at the floor of the film's Watch on from the first draft (44 px for desktop); the film's first draft had about 40 at 24–40 px. The line used as a shot's title counts as a 主标题. Source, credit and asset lines go into the video description unless a licence or the type doc wants them on screen (a data story's source line), and then they are labels at the floor.
  • playbook/02-verification.md layer 3: a strip at each text swap in one place (swatch: render on the CPU and encode single-threaded so every run gives the same bytes #17's old-out-before-new-in; the reviewer found three overlaps the author's own check missed) and at each rolling readout, whose resting values are checked against NOTES.
  • engines/README.md (HyperFrames + Three.js pitfalls): RoomEnvironment reflections drew a bright vertical strip down every glass flask, and a full-turn LatheGeometry seam faces a +z camera by default (phiStart = Math.PI turns it away). The glass scene drops scene.environment (the other materials there lose it too; light them instead) and gets a fresnel rim shell: the same geometry scaled about 1.004, additive ShaderMaterial, no depth write, pow(1 − |N·V|, 2.6). The list's lead-in now says "most" of its pitfalls come from the intro film (Lessons from four films: copy that reads as written by a person, render and audio pitfalls, contributing notes #66 had already added others).
  • CHANGELOG.md: an Unreleased entry.

Floors touched (CONTRIBUTING rule 7)

Not in this PR

  • A command for the pace evening, and a whisper path for word times in bin/vh tts: each its own PR if wanted.
  • tools/audio/tts.py decides "daily quota" only from the delay a 429 asks for; the quota id in the error body (…PerDay…) would say so directly. Reported separately, not changed here.

Verified

  • Against tools/audio/tts.py: in a joined gemini request every line is a text item with a speech_metadata annotation whose style is the overall style joined with that line's direction; the per-line path writes the take before the first transcription call, and --join writes _blockNN.wav before transcribing it; a 429 asking for 90 s or less (60 s when it names no delay) is retried up to 5 times a call; align_lines spreads script units over a heard unit they do not match one to one ("一九五三" over "1953").
  • The pace numbers were re-measured: the film's script re-run on a scratch copy of its raw take gives 4.71–6.67 chars/s per line after the cuts and the pause cap (sd 0.45), 5.56–6.13 after (mean 5.92, sd 0.11). The film's notes said 5.0–7.8 before, a different measure, so the docs use these ranges (and say they are a reference, not a target).
  • Against three.js 0.181.2's LatheGeometry: vertex x = r·sin φ, z = r·cos φ, so the φ = 0 seam faces +z.
  • tools/ci.sh --committed passes, also with VH_BASH=/bin/bash; shellcheck and pyflakes ran through the uv recipe. No keys, private paths or film names in the diff.

Independent review (one Sonnet reviewer, read-only, fresh context): the code facts held. Fixed from it: the source-line advice contradicted type 05's on-screen source line (now: on screen when a licence or the type doc wants it, as a label at the floor); the year/whisper mechanism (the harness aligner does not start late; a home-made one can) and a pointer to align_lines; the one-voice advice limited to single-narrator films and to one short style; the pace numbers, the factor formula, the 8-bit padding pitfall and writing the timeline and captions back; the two floor changes stated here and in the CHANGELOG, with #6 saying the headline rule itself; the README's quota text describing what the code does; small wording.

ZLHad added 3 commits October 7, 2026 21:48
…, labels at the floor from the start, text swaps and readouts, glass in Three.js

The maintainer asked that the general lessons of a 2-minute narrated canvas + Three.js film go into the docs as short guidance.

- playbook/04: in a one-request gemini take every line still carries its own style, so a per-line [direction] makes that line sound different; for one voice, no per-line directions. Evening out the pace line by line on one take (waveform-refined cuts, capped inner pauses, partial per-line atempo towards one target), described, not a command. When transcription is unavailable: the take is on disk; --resume later or align locally with whisper, mapping digits back to spelled-out years.
- tools/audio/README: a daily-quota 429 was once waited on and retried like the per-minute limit, which looked hung.
- playbook/03 §4: write labels at the floor from the first draft; a shot's headline is a 主标题; credits go into the description.
- playbook/02 layer 3, TASTE_CHECKLIST #19: strips at text swaps and rolling readouts; a readout rests only on sourced values.
- engines/README: RoomEnvironment reflections and the default LatheGeometry seam on glass; a fresnel rim shell instead.
- playbook/04: whisper output goes through tts.py's align_lines (it spreads spelled-out years over a heard "1953"; a home-made character aligner can misplace the line start); the one-voice advice is for single-narrator films and asks for one short overall style; the pace numbers are re-measured on the take with the same method (4.7–6.7 → 5.6–6.1 chars/s, sd 0.45 → 0.11); the evening steps say how the factor is computed, avoid anullsrc + concat, and write the take, the timeline and the captions back.
- tools/audio/README: the quota text describes what the code does (a 429 under 90 s, or without a delay, is retried up to 5 times) and points to the playbook; the joined-synthesis section points to the one-voice advice.
- playbook/03 §4: credits stay on screen where a licence or the type doc wants them (a data story's source line), as labels at the floor; the headline is "the line used as the shot's title".
- TASTE_CHECKLIST #6: a shot's headline counts as a 主标题.
- engines/README: "a row of glass objects"; dropping scene.environment also drops it for the other materials in that scene; the list's lead-in says most pitfalls come from the intro film.
- playbook/02: "旧句退完、新句才进", as in #17.
- CHANGELOG: rewritten to match, with the two floor changes said explicitly.
…umber

The maintainer: "别太严格这个语速". The evening is optional (nothing to do when it sounds fine), aims for closer rather than equal, and the film's numbers are a reference.
@ZLHad
ZLHad marked this pull request as ready for review October 7, 2026 14:30
@ZLHad
ZLHad merged commit eca390d into main Oct 7, 2026
2 checks passed
@ZLHad
ZLHad deleted the claude/lessons-one-take-voice-labels branch October 7, 2026 14:30
ZLHad added a commit that referenced this pull request Oct 7, 2026
The ci.sh stub check sets no_proxy for 127.0.0.1, so an exported http_proxy
cannot turn it red, and covers the fallback too: a 429 naming no daily quota
that asks for 120 s fails at once. tools/audio/README says again that a 429
with no delay counts as 60 s; the #70 changelog entry points at the new one.
ZLHad added a commit that referenced this pull request Oct 7, 2026
* bin/vh tts: a 429 that names a per-day Gemini quota fails at once

Gemini's 429 for a spent daily quota can ask for a retryDelay under a
minute, and gemini_call took a 429 for the daily quota only when it asked
for more than 90 s. Below that it waited and retried up to 5 times a call,
so --align gemini looked hung once the daily quota was spent.

gemini_call now reads the quotas the 429 body names: a google.rpc.QuotaFailure
violation whose quotaId or quotaMetric has "PerDay" in it fails the call at
once with a message naming the quota and the --resume recovery. Per-minute
429s are waited out and retried as before; a 429 naming no daily quota still
fails at once only when it asks for more than 90 s.

tools/ci.sh checks both cases against a stub server on 127.0.0.1. The quota
sentences in tools/audio/README and playbook/04 say the daily quota now
fails at once.

* Fixes from the independent review

The ci.sh stub check sets no_proxy for 127.0.0.1, so an exported http_proxy
cannot turn it red, and covers the fallback too: a 429 naming no daily quota
that asks for 120 s fails at once. tools/audio/README says again that a 429
with no delay counts as 60 s; the #70 changelog entry points at the new one.
ZLHad added a commit that referenced this pull request Oct 7, 2026
…, without synthesizing again (#72)

A one-request take can run audibly fast and slow from line to line. playbook/04 described how to even it out on the
take (#70); this makes it a command. It stays gentle, as the maintainer asked ("别太严格这个语速"): optional, each line
moves part of the way towards the median (or --ref / --target), lines already close keep their pace, and --dry-run
shows the numbers before anything is written.

Lines split at the silence nearest their ASR boundary; cut points refined from the waveform; long pauses inside a line
capped unless the script asked for them; ffmpeg atempo per line; the take's own pauses between lines kept inside a
range by punctuation and block. Word times carried through the edit by default (or transcribed again with gemini or a
local whisper, through tts.py's align_lines). The originals stay as .raw; a re-run starts from them, --restore puts
them back. tools/ci.sh checks it on a synthetic take with known syllable onsets.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant