Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 7 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,10 +2,16 @@

## Unreleased

**`bin/vh tts --align gemini`: a spent daily quota fails at once instead of looking hung**
- Why: Gemini's 429 for a spent per-day quota can ask for a short `retryDelay` (27 s, 11.8 s in reports on Google's forum), and tts.py took a 429 for the daily quota only when it asked for more than 90 s. Below that it waited and retried up to 5 times a call, 4 calls at a time, so a 2026-10 film's `--align gemini` looked hung (documented in the entry below).
- tts.py reads the quotas a 429 names: when a `google.rpc.QuotaFailure` violation's `quotaId` or `quotaMetric` has `PerDay` in it (`GenerateRequestsPerDayPerProjectPerModel`, also when it is listed with the per-minute quota), the call fails at once, whatever delay it asks for, with a message that names the quota and says `--resume` continues the run once it resets. The take is kept as before. Per-minute 429s are waited out and retried as before, and a 429 that names no daily quota still fails at once only when it asks for more than 90 s. The same applies to `bin/vh voices`.
- Reproduced with a local stub server returning a per-day 429 with `retryDelay` 30s on a one-line `say` script: before, 5 waits of 31 s (157 s) and 6 requests per line, then the failure; after, one request and the failure at once, with `voiceover.zh.wav`, the timeline (the line keeps its timing, `asr.error` names the quota) and `_run.json` on disk. A per-minute 429 is still retried 5 times.
- tools/ci.sh: a stub on 127.0.0.1 checks the per-day, per-minute and over-90 s cases (fails on the old tts.py). tools/audio/README ("转写的配额" and the limits paragraph) and playbook/04 ("转写用不了时") say the daily quota fails at once.

**Lessons from a narrated film with 3D: one voice across a one-request take, labels sized from the start, text swaps and readouts, glass in Three.js**
- Why: the maintainer asked that the general lessons from a 2-minute narrated canvas + Three.js film go into the docs as short guidance.
- playbook/04: in a one-request gemini take (`--join all`) every line still carries its own style, so one line with a `[direction]` (often the opening) tends to sound different from the rest; for one voice in a single-narrator film, write one short overall style and no per-line directions (tools/audio/README's "整段合成" points there). When lines in one take run audibly fast and slow, an optional way to bring them closer line by line without synthesizing again, not to one exact pace (cuts refined from the waveform, so a first syllable the ASR placed late is not clipped; long pauses inside a line capped; per-line `atempo` part of the way towards one target pace; the take and the timeline written back, aligned again, captions redone), described, not a command yet: the film went from 4.7–6.7 to 5.6–6.1 CJK chars/s per line; the numbers are a reference, the ear decides. When transcription is unavailable (daily quota, API errors): the take is already on disk; `--resume` later, or take word times from a local whisper and align them with tts.py's `align_lines`, mapping digits back to a script that spells years out (一九五三).
- tools/audio/README: a 429 asking for under 90 s (or naming no delay) is retried like the per-minute limit, up to 5 times a call, so a spent daily quota can look hung; the old text said the daily quota always fails at once.
- tools/audio/README: a 429 asking for under 90 s (or naming no delay) is retried like the per-minute limit, up to 5 times a call, so a spent daily quota can look hung; the old text said the daily quota always fails at once. (Superseded by the entry above: a 429 that names the daily quota now fails at once.)
- playbook/03 §4: write readable labels at the floor from the first draft (44 px for `desktop`; one draft had about 40 at 24–40 px). Source, credit and asset lines go into the description unless a licence or the type doc wants them on screen, and then they are labels.
- playbook/02 layer 3: a strip at each text swap in one place (#17's old-out-before-new-in, unchanged; a reviewer found three that the author's own check missed) and at each rolling readout, whose resting values are checked against NOTES.
- Floors, said explicitly (CONTRIBUTING rule 7): TASTE_CHECKLIST #19 names a readout resting on an interpolated value as a fabricated number (the film's reviewer had judged it under #19); #6 and playbook/03 §4 count a shot's one-line headline as a 主标题, so it is held to 84 px (150 for `feed`), not 44.
Expand Down
2 changes: 1 addition & 1 deletion playbook/04-audio.md
Original file line number Diff line number Diff line change
Expand Up @@ -70,7 +70,7 @@ bin/vh beats <任意音乐文件> # 外来音乐的节

`--align gemini` 对任何 provider 都能用:每句合成完交给 Gemini 转写,拿回带时间戳的词,再和稿子逐字对齐。结果写进 timeline:`words` 是逐词时间(文字用稿子的写法),字幕和画面逐词同步用它;`asr` 是和稿子的相似度,跑偏、或者句首句尾多出 1 s 以上声音的句子标上 `flag`,命令最后逐句列出来。转写失败也不丢配音:修好原因后,同一条命令加 `--resume` 接着跑。时间步长、费用、限额和实测见 `tools/audio/README.md` 的同名一节。

**转写用不了时**(当天额度用完,或者接口一直报错):额度用完后,`--align gemini` 可能在 429 上等了又试(每次调用最多重试 5 次),看着像卡住。合成好的音频已经在盘上(逐句合成时是 `voiceover.<lang>.wav`,`--join` 时是 `vo/<lang>/_blockNN.wav`),停掉不会丢;可以等额度恢复后加 `--resume` 接着转写,也可以改用本地 whisper 拿词级时间(还没封装;Apple Silicon 上用 mlx-whisper 的 `whisper-large-v3-turbo`,`initial_prompt` 里写上人名和术语;别的做法见下文"选型"的词级时间戳一行),再交给 `tools/audio/tts.py` 的 `align_lines` 和稿子对齐(`--align gemini` 用的也是它)。whisper 常把年份写成阿拉伯数字,和稿子里为念法写的"一九五三"对不上:相似度会掉,自己写的逐字对齐还可能把那句的起点算偏(一支 2026-10 的片子算晚了,剪掉了第一个字)。对齐前先把这些数字换回稿子的写法。
**转写用不了时**(当天额度用完,或者接口一直报错):额度用完后,`--align gemini` 直接报错退出,报错里写出是哪一项日配额(429 没写明配额、又只要求等几十秒时,仍会等了再试,每次调用最多 5 次)。合成好的音频已经在盘上(逐句合成时是 `voiceover.<lang>.wav`,`--join` 时是 `vo/<lang>/_blockNN.wav`),停掉不会丢;可以等额度恢复后加 `--resume` 接着转写,也可以改用本地 whisper 拿词级时间(还没封装;Apple Silicon 上用 mlx-whisper 的 `whisper-large-v3-turbo`,`initial_prompt` 里写上人名和术语;别的做法见下文"选型"的词级时间戳一行),再交给 `tools/audio/tts.py` 的 `align_lines` 和稿子对齐(`--align gemini` 用的也是它)。whisper 常把年份写成阿拉伯数字,和稿子里为念法写的"一九五三"对不上:相似度会掉,自己写的逐字对齐还可能把那句的起点算偏(一支 2026-10 的片子算晚了,剪掉了第一个字)。对齐前先把这些数字换回稿子的写法。

有词级时间时,`bin/vh captions` 会给每条字幕加上 `words`,念得太快的短字幕会往后延到 1.8 s(`tools/audio/README.md` 的"字幕")。

Expand Down
4 changes: 2 additions & 2 deletions tools/audio/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@
- 模拟 Qwen 停不下来:同一句后面接 8 s 低音量的无关絮语,相似度掉到 0.47,并报"最后一个词之后还有 5.5 s";接 8 s 近乎无声的底噪,文字全对(1.00),但报"最后一个词之后还有 8.2 s"。两种都抓到了;
- 双人对话里,ASR 会把 B 说话时 A 的一声"嗯"也转出来。这种插话按稿子里的 `|嗯|` 算,不扣分。

局限:时间戳的步长是 0.1 s(30 fps 下 3 帧);数字可能被规范化("二十六"转成"26"),相似度会降,但不算念错;每次调用有 3–7 s 延迟,逐句模式下 4 句并发。费用约 $0.005 / 分钟音频(付费档),免费档不收费。Tier 1 每分钟只能转写 10 次,超过 10 句(算上复查)时会碰到 429:这时按接口给的时间等(通常不到 1 分钟)再重试,12 句实测 155 s 跑完。接口要求等 90 s 以上的当成日配额用完,直接报错(日配额也可能只要求等几十秒,见下文"转写的配额")。转写失败到底(日配额、断网、某个文件坏了)也不丢配音:`voiceover.<lang>.wav` 和 timeline 在合成后就先写好,出错的句子保留实测时长、`asr` 里记 `error`,命令以非 0 退出;修好原因后用**同一条命令**加 `--resume` 接着跑:只补合成缺的文件(合成中途失败也一样),转写成功过的句子不再重转(结果存在 `vo/<lang>/*.asr.json`),其余重做;`--join` 时整段音频留在 `vo/<lang>/_blockNN.wav`,在那之前它的几句共用整段的起止(`minimax` 例外:句子照样按它自己的逐字时间切开)。那次运行的稿子、provider、音色、语气指示(`--instruct`)和 `--join` 方式记录在 `vo/<lang>/_run.json` 里,`--resume` 要求它们一样,改了就拒绝并指出改了什么(对拍、间距这些时间参数可以改)。不带 `--align` 的普通合成中途失败,也用同一条命令加 `--resume` 接着合成;合成好的普通配音之后再加 `--align gemini --resume`,只做转写,不再合成。key 在合成前就检查。
局限:时间戳的步长是 0.1 s(30 fps 下 3 帧);数字可能被规范化("二十六"转成"26"),相似度会降,但不算念错;每次调用有 3–7 s 延迟,逐句模式下 4 句并发。费用约 $0.005 / 分钟音频(付费档),免费档不收费。Tier 1 每分钟只能转写 10 次,超过 10 句(算上复查)时会碰到 429:这时按接口给的时间等(通常不到 1 分钟)再重试,12 句实测 155 s 跑完。429 写明是日配额用完的(见下文"转写的配额"),或者要求等 90 s 以上的,直接报错,不再等。转写失败到底(日配额、断网、某个文件坏了)也不丢配音:`voiceover.<lang>.wav` 和 timeline 在合成后就先写好,出错的句子保留实测时长、`asr` 里记 `error`,命令以非 0 退出;修好原因后用**同一条命令**加 `--resume` 接着跑:只补合成缺的文件(合成中途失败也一样),转写成功过的句子不再重转(结果存在 `vo/<lang>/*.asr.json`),其余重做;`--join` 时整段音频留在 `vo/<lang>/_blockNN.wav`,在那之前它的几句共用整段的起止(`minimax` 例外:句子照样按它自己的逐字时间切开)。那次运行的稿子、provider、音色、语气指示(`--instruct`)和 `--join` 方式记录在 `vo/<lang>/_run.json` 里,`--resume` 要求它们一样,改了就拒绝并指出改了什么(对拍、间距这些时间参数可以改)。不带 `--align` 的普通合成中途失败,也用同一条命令加 `--resume` 接着合成;合成好的普通配音之后再加 `--align gemini --resume`,只做转写,不再合成。key 在合成前就检查。


需要更细的强制对齐(音素级、离线)时,仍然可以用 mlx-audio 的 Qwen3-ForcedAligner、FunASR 或 whisper.cpp,这些没有封装。
Expand Down Expand Up @@ -95,7 +95,7 @@

- **费用**(Standard 档,每百万 token):Flash TTS 输入 $0.50、输出音频 $9(约 $0.00225 / 10 s);Flash-Lite TTS 输入 $0.50、输出 $6(约 $0.0015 / 10 s)。这是 2026-12-31 之前的价格,**2027-01-01 起全部翻倍**(Flash $1 / $18,Lite $1 / $12)。Batch 和 Flex 档是一半。免费档不收费,但内容可能被用来改进 Google 的产品,有保密要求的稿子用付费档;
- 单个请求最多 8,192 个输入 token;
- **转写的配额**(`--align gemini`,2026-10-01 在 Tier 1 上):`gemini-3.5-transcribe` 每分钟 10 次、每天 100 次,官方文档说按项目算,不按 key。每分钟的 429 会自动等;每天的用完时,429 要求等 90 s 以上就直接失败;要求不到 90 s(或没写,按 60 s 算)时,会像每分钟的限额一样等了再试,每次调用最多重试 5 次。有一支片子在日配额用完后就是这样,看着像卡住,怎么办见 `playbook/04-audio.md` 的"转写用不了时"。额度恢复后用 `--resume` 接着跑,已经转写成功的句子(`vo/<lang>/*.asr.json`)不会重做。每天几点恢复没有定论:文档写太平洋时间零点,429 提示里的倒计时指向 UTC 零点。AI Studio 的用量页显示的是 28 天的峰值,不是今天还剩多少;
- **转写的配额**(`--align gemini`,2026-10-01 在 Tier 1 上):`gemini-3.5-transcribe` 每分钟 10 次、每天 100 次,官方文档说按项目算,不按 key。每分钟的 429 会自动等;每天的用完了会直接失败,报错里写出是哪一项日配额。日配额的 429 也可能只要求等几十秒,所以不看等多久,看 429 里列出的配额:`google.rpc.QuotaFailure` 的 `quotaId` 带 `PerDay`(例如 `GenerateRequestsPerDayPerProjectPerModel`)就是日配额。没写明配额的 429,要求等 90 s 以上才当成日配额,不到 90 s(或没写,按 60 s 算)的照样等了再试,每次调用最多 5 次。转写用不了时怎么办,见 `playbook/04-audio.md` 的"转写用不了时"。额度恢复后用 `--resume` 接着跑,已经转写成功的句子(`vo/<lang>/*.asr.json`)不会重做。每天几点恢复没有定论:文档写太平洋时间零点,429 提示里的倒计时指向 UTC 零点。AI Studio 的用量页显示的是 28 天的峰值,不是今天还剩多少;
- **语气写短**:`--instruct` 和 `[指示]` 最后都进 `speech_metadata.style`。官方建议只写这一句的情境语气("压低声音,卖个关子");年龄、性别、口音这类身份特征不要写进 style,要换音色,或者用 `voices design` 做一个。长篇的"角色设定""导演笔记"是音色漂移最常见的原因;
- **每段 Gemini 音频都带 SynthID 水印**(听不出来,但能检测到)。用了 Gemini 旁白的片子,在项目 `NOTES.md` 里写明"旁白为 AI 合成(Gemini TTS,含 SynthID 水印)",发布时按平台要求标注。

Expand Down
22 changes: 17 additions & 5 deletions tools/audio/tts.py
Original file line number Diff line number Diff line change
Expand Up @@ -138,8 +138,8 @@ def to_wav(src: Path, dst: Path):
def duration(p: Path) -> float:
return float(run(["ffprobe", "-v", "error", "-show_entries", "format=duration", "-of", "csv=p=0", str(p)]).stdout)

RATE_WAIT_MAX, RATE_TRIES = 90.0, 5 # a 429 asking for more than 90 s is a daily quota: waiting will not help
BODY_PARSE, BODY_SHOW = 8192, 500 # bytes of an error body read for the retry delay / repeated in the message
RATE_WAIT_MAX, RATE_TRIES = 90.0, 5 # a 429 asking for more than 90 s is taken for a daily quota: waiting will not help
BODY_PARSE, BODY_SHOW = 8192, 500 # bytes of an error body read for the retry delay and the quota / repeated in the message

class GeminiError(RuntimeError):
"""A Gemini call that failed for good (after its retries). tts keeps what it has and reports it; voices exits."""
Expand All @@ -153,13 +153,21 @@ def retry_after(e, msg):
except ValueError: # Retry-After as an HTTP date
return 61.0

def daily_quota(body):
"""The per-day quota a 429 body names, else None: the quotaId or quotaMetric of a google.rpc.QuotaFailure violation
with "PerDay" in it (GenerateRequestsPerDayPerProjectPerModel). One 429 can list the per-minute and the per-day
quotas together, and a spent daily quota can ask for a retryDelay under a minute, so the delay cannot tell them apart."""
m = re.search(r'"(?:quotaId|quotaMetric)":\s*"([^"]*PerDay[^"]*)"', body)
return m and m.group(1)

def gemini_call(path, body=None, method=None, query=None, timeout=180, tries=3):
"""One Gemini REST call. The key travels in a header only, never in the URL or an error message.
5xx and network errors are retried twice: the TTS docs note rare 500s when the model returns text instead of audio.
A 429 waits as long as the API asks and is retried up to 5 times: Tier 1 allows 10 transcribe calls a minute, so
--align on 11 or more lines hits it. A 429 asking for more than 90 s (a daily quota) fails at once.
The delay is read from the first 8 KB of the body: a long quota message puts "retry in Ns" past the first 500 bytes,
which are all that the error message repeats. A call that fails for good raises GeminiError."""
--align on 11 or more lines hits it. A 429 that names a per-day quota fails at once, whatever delay it asks for, and
so does one asking for more than 90 s. The delay and the quota are read from the first 8 KB of the body: a long quota
message puts "retry in Ns" past the first 500 bytes, which are all that the error message repeats. A call that fails
for good raises GeminiError."""
key = os.environ.get("GEMINI_API_KEY") or sys.exit("set GEMINI_API_KEY (Google AI Studio → Get API key)")
url = f"{GEMINI_API}/{path}" + (f"?{urllib.parse.urlencode(query, doseq=True)}" if query else "")
data = None if body is None else json.dumps(body).encode()
Expand All @@ -173,6 +181,10 @@ def gemini_call(path, body=None, method=None, query=None, timeout=180, tries=3):
return json.loads(raw) if raw.strip() else {}
except urllib.error.HTTPError as e:
body = e.read()[:BODY_PARSE].decode(errors="replace")
quota = daily_quota(body) if e.code == 429 else None
if quota:
raise GeminiError(f"gemini {path.split('/')[0]}: the daily quota is spent ({quota}); retrying will not help"
f" until it resets (bin/vh tts: then --resume continues the run)")
wait = retry_after(e, body) if e.code == 429 else None
if wait is not None and wait <= RATE_WAIT_MAX and limited < RATE_TRIES:
limited += 1
Expand Down
Loading
Loading