Skip to content

bin/vh tts: a 429 that names a per-day Gemini quota fails at once - #71

Merged
ZLHad merged 2 commits into
mainfrom
claude/gemini-daily-quota
Oct 7, 2026
Merged

ZLHad merged 2 commits into
mainfrom
claude/gemini-daily-quota

Conversation

@ZLHad

@ZLHad ZLHad commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

Why

gemini_call in tools/audio/tts.py decided "daily quota" only from the delay a 429 asks for: over 90 s (RATE_WAIT_MAX) failed at once, anything shorter (or no delay, taken as 60 s) was waited out and retried up to 5 times a call (RATE_TRIES). Gemini's 429 for a spent per-day quota can ask for a short retryDelay, so once the daily quota was spent, bin/vh tts … --align gemini kept waiting and retrying and looked hung (a 2026-10 film; documented as a known behaviour in tools/audio/README "转写的配额" and playbook/04 "转写用不了时" by #70).

What the real 429s look like (checked, not guessed):

  • google/rpc/error_details.proto: QuotaFailure { repeated Violation violations }, with Violation fields subject, description, api_service, quota_metric, quota_id, quota_dimensions, quota_value, future_quota_value; Gemini's REST bodies spell them quotaMetric, quotaId, … and add a google.rpc.RetryInfo with retryDelay.
  • Reported Gemini 429s: quotaId: "GenerateRequestsPerDayPerProjectPerModel-FreeTier" with retryDelay: "27s"; another listed …PerMinute… and …PerDay… violations together with retryDelay: "11.767792953s" (Generate error: 429 RESOURCE_EXHAUSTED. google-gemini/gemini-cli#8437 has the same mix with "55s"). In these the quotaMetric is e.g. generativelanguage.googleapis.com/generate_content_free_tier_requests, so in practice the quotaId is what names the day.

What changed

  • tools/audio/tts.py: new daily_quota(body) returns the quotaId / quotaMetric of a QuotaFailure violation with PerDay in it, read from the first 8 KB of the body that gemini_call already reads. A 429 that names one raises GeminiError at once, whatever delay it asks for: gemini interactions: the daily quota is spent (GenerateRequestsPerDayPerProjectPerModel); retrying will not help until it resets (bin/vh tts: then --resume continues the run). Per-minute 429s are unchanged (waited out, up to 5 retries), and a 429 naming no daily quota still fails at once only above 90 s. bin/vh voices goes through the same call.
  • tools/ci.sh (smoke): a stub server on 127.0.0.1, GEMINI_API pointed at it, time.sleep recorded rather than slept: a 429 listing the per-minute and per-day quotas with retryDelay 30s must raise after one request and no wait, naming the quota and --resume; two per-minute 429s and then a 200 must return the 200 after two 31 s waits; a 429 naming no daily quota that asks for 120 s must fail at once (the old fallback). no_proxy is set for 127.0.0.1 so an exported http_proxy cannot turn it red. It fails on the old tts.py ("a per-day 429 was retried (waits [31.0])"). Standard library only, under a second.
  • Docs: tools/audio/README ("转写的配额" and the limits paragraph of the alignment check, which repeated the 90 s rule) and playbook/04 ("转写用不了时") say the daily quota fails at once and name the error; CHANGELOG.md Unreleased.

Before / after

Reproduced with a local stub that answers every call with a per-day 429 (retryDelay 30s, body shaped like the real ones above), on a one-line say script with --align gemini, running tts.py's main() in-process with GEMINI_API pointed at the stub and the sleeps recorded:

requests to the stub waits asked exit on disk
before 6 5 × 31.4 s = 157 s 1 voiceover.zh.wav, timeline (asr.error = the raw 429 JSON), _run.json
after 1 none 1 the same; asr.error names GenerateRequestsPerDayPerProjectPerModel

After, the run ends with:

align: 1 line(s) not transcribed (hook): gemini interactions: the daily quota is spent (GenerateRequestsPerDayPerProjectPerModel); retrying will not help until it resets (bin/vh tts: then --resume continues the run). The take is written (voiceover and timeline above); fix the cause, then re-run the same command with --resume to redo only what failed

The same stub with only the per-minute quota in the body gives 5 waits of 31.4 s and then the failure both before and after: per-minute handling is unchanged.

Before this fix, with 4 lines transcribed at a time, each call spent about 2.5 minutes in waits before failing, so a long script sat in waits for many minutes after the daily quota ran out.

How it was verified

  • The stub reproduction above, before and after.
  • tools/ci.sh --committed with VH_BASH=/bin/bash (bash 3.2) and with the default bash 5, shellcheck and pyflakes through the uv shim from CONTRIBUTING: 104 checks pass, none skipped.
  • No live Gemini call was made (the stub covers both cases); git grep finds no AIzaSy… or AQ.… key.
  • An independent review (Sonnet) found no blocker; it re-ran the CI snippet against old and new tts.py (20 green runs on the new one) and the end-to-end stub with a 2-line script. Fixed from it: the stub check under an exported http_proxy, the over-90 s fallback case added to the check, the "no delay counts as 60 s" detail back in tools/audio/README, and a pointer from the Lessons from a narrated film with 3D: one voice in a one-request take, labels at the floor from the start, text swaps and readouts, glass in Three.js #70 changelog entry to this one. Kept on purpose: the --resume hint in the GeminiError message (asked for; bin/vh voices sees it prefixed bin/vh tts:), and quotaMetric in the match (real metrics are snake_case, so quotaId is what matches in practice).
  • Not covered: whether a spent daily quota on other Gemini models or tiers always names PerDay in quotaId. One that does not still falls back to the old 90 s rule.

ZLHad added 2 commits October 7, 2026 22:36
Gemini's 429 for a spent daily quota can ask for a retryDelay under a
minute, and gemini_call took a 429 for the daily quota only when it asked
for more than 90 s. Below that it waited and retried up to 5 times a call,
so --align gemini looked hung once the daily quota was spent.

gemini_call now reads the quotas the 429 body names: a google.rpc.QuotaFailure
violation whose quotaId or quotaMetric has "PerDay" in it fails the call at
once with a message naming the quota and the --resume recovery. Per-minute
429s are waited out and retried as before; a 429 naming no daily quota still
fails at once only when it asks for more than 90 s.

tools/ci.sh checks both cases against a stub server on 127.0.0.1. The quota
sentences in tools/audio/README and playbook/04 say the daily quota now
fails at once.
The ci.sh stub check sets no_proxy for 127.0.0.1, so an exported http_proxy
cannot turn it red, and covers the fallback too: a 429 naming no daily quota
that asks for 120 s fails at once. tools/audio/README says again that a 429
with no delay counts as 60 s; the #70 changelog entry points at the new one.
@ZLHad
ZLHad marked this pull request as ready for review October 7, 2026 15:17
@ZLHad
ZLHad enabled auto-merge (squash) October 7, 2026 15:17
@ZLHad
ZLHad merged commit 5726037 into main Oct 7, 2026
2 checks passed
@ZLHad
ZLHad deleted the claude/gemini-daily-quota branch October 7, 2026 15:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant