Skip to content

bin/vh tts: a minimax provider (MiniMax T2A v2), timed by its own subtitles - #69

Merged
ZLHad merged 2 commits into
mainfrom
claude/minimax-tts
Oct 7, 2026
Merged

ZLHad merged 2 commits into
mainfrom
claude/minimax-tts

Conversation

@ZLHad

@ZLHad ZLHad commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

Why

A film's narration (2026-10-07) was made with MiniMax T2A speech-2.8-hd through a project script. Two things made it worth a provider:

  • the whole script went out in one request, so the voice could not drift between blocks (the maintainer had heard and disliked that drift with block-by-block synthesis);
  • the request's own subtitles (subtitle_type: word) gave every character's time, so no ASR pass was needed for line boundaries or word captions.

What changed

tools/audio/tts.py: new provider minimax

  • POST https://api.minimax.io/v1/t2a_v2 (the China host api.minimaxi.com rejects international keys with 2049; MINIMAX_BASE_URL for a China-platform key).
  • Key from MINIMAX_TOKEN_PLAN_KEY (Token Plan, spends its credits) first, then MINIMAX_API_KEY (pay-as-you-go). Errors name the variable, never the key, with a hint for 1008 (no balance on that billing) and 1004/2049 (key and host don't match). Checked before anything is deleted.
  • MINIMAX_TTS_MODEL (default speech-2.8-hd), MINIMAX_TTS_SPEED / _VOL / _PITCH: range-checked up front and recorded in vo/<lang>/_run.json, so --resume refuses a change. text_normalization on, language_boost from --lang, 44.1 kHz WAV converted to 48 kHz like every provider.
  • --instruct / [direction] count only when one of MiniMax's nine emotion words; anything else is ignored, said once.
  • Word timing: subtitle entries are grouped back into the words they were read from ("1953" spans the seven syllables of 一千九百五十三, a @pronounce name its pinyin, "Claude" comes as C / la / ude) and found in order in the text sent, then cut into the same script-spelled units elevenlabs gives. The file's character offsets are not used: one response counted the pause marks in them, another did not.
  • --join block|all with minimax places the lines by those words: no --align gemini, no GEMINI_API_KEY. --align gemini still works (one transcription per block for the similarity check; times stay MiniMax's). Words are kept beside the audio (NN.words.json, _blockNN.words.json); --resume treats a line or block as done only with them.
  • Pause marks <#0.4#>: voiced by minimax; captions and every other provider (gemini included) drop them. In a joined request lines are linked by <#--gap#> plus marks written at line edges; marks at a request's edges are dropped and runs merged, as MiniMax requires.
  • @pronounce 硖合/(xia2)(he2) lines in script.txt → pronunciation_dict.tone (minimax only; other providers print a note). A changed dictionary makes --resume refuse.

Docs: tools/audio/README.md (provider table; a "MiniMax" section: keys and hosts, voices and how to list them via POST /v1/get_voice, settings, emotion words, word timing, one request for the whole script, pause marks, @pronounce, billing, measurements, and the year pitfall: "1953 年" is read 一千九百五十三年, so write 一九五三年; 2001/2003 read 两千零一/两千零三, acceptable); playbook/04-audio.md (commands, provider table, who can act, tags, 写念法, selection table); bin/vh help and doctor; type 02; captions.py docstring; CHANGELOG.md.

Known limitation (documented): MiniMax's Chinese voice ids contain a space, so they can't go into @speakers; a minimax dialogue isn't possible yet. The default stays --join none like every other provider; the docs' examples use --join all.

How it was verified

Against the live API (Token Plan key, 2026-10-07):

  • --join all (3 zh lines, 1 request, ~9 s): line spans and per-character words from the subtitles, @pronounce sent, an inline <#0.4#> gave ~0.8 s with the comma; no Gemini call.
  • Per line with calm; 平静地讲 (emotion used, the rest noted once), --join block with a non-default voice and MINIMAX_TTS_SPEED=1.2, English side with happy (3.5 and GHz. timed separately).
  • --resume: a finished run re-synthesizes nothing; a deleted block, or a block whose words file is missing, is synthesized again (the other block's audio unchanged, same md5); a changed speed or @pronounce entry is refused, naming what changed.
  • Errors: pay-as-you-go key only → 1008 with the hint; MINIMAX_BASE_URL=https://api.minimaxi.com → 2049 with the host hint; no key and a speed out of range exit before anything is deleted.
  • --join all --align gemini: Gemini's daily transcription quota was used up, which exercised the failure path (lines still cut and timed by MiniMax, asr: {error}, take written, exit 1, no "block kept whole" message). The success path was run from a cached .asr.json with --resume: per-line similarity and flag, head/tail at the block edges, times unchanged.
  • bin/vh captions on the result: captions.json items carry words, .lines.srt written.

Other providers unchanged (old tts.py from main vs new, same script with tags, reactions, bilingual lines):

  • say per line, zh and en: timeline, voiceover and _run.json byte-identical (one first en run of say differed from every later run, old or new; repeated old runs match the new one);
  • say --join block from the same cached transcription: timeline, voiceover and every cut file byte-identical; with transcription failing (429): same output and timeline.

tools/ci.sh --committed passes, also with VH_BASH=/bin/bash (shellcheck and pyflakes through the uv shim).

Independent review

One reviewer (fresh context) re-ran the old-vs-new comparisons with stubs and found no blocking bug. Fixed in the second commit, each re-tested against the API or offline:

  • a pause mark at a line's edge was dropped between requests (per line, --join block), though the README's example relied on it: it now adds to the silence there too (per line: exactly --gap + mark between the two lines);
  • outside a dialogue, a spoken |x| was missing from words;
  • the 10,000-character limit was checked only after the last take had been deleted: now before (the old files stay);
  • --gap is part of a joined minimax request: recorded in _run.json, --resume refuses a change;
  • a removed pause mark put a space between a CJK character and a Latin word; one spacing rule now serves captions and the request;
  • @pronounce without a / stays a narration line with that id, as before;
  • docs: volume range, the joined-run failure note, CHANGELOG wording.
    Not changed: a near-limit --join all returns its WAV as hex in the JSON (about 80 MB for a 457 s film, fine at our lengths).

After the fixes, say per line (zh, en, --beats --snap downbeat, --gap/--lead) and --join block with --beats are again byte-identical old vs new; tools/ci.sh --committed passes (also VH_BASH=/bin/bash).

ZLHad added 2 commits October 7, 2026 19:49
…titles

A film's narration was made with MiniMax speech-2.8-hd through a project
script: the whole script in one request kept the voice from drifting
between blocks, and the request's own subtitles gave every character's
time without an ASR pass. This makes it a provider.

- POST https://api.minimax.io/v1/t2a_v2; key from MINIMAX_TOKEN_PLAN_KEY,
  then MINIMAX_API_KEY (never printed); hints for 1008 and 1004/2049;
  MINIMAX_BASE_URL for a China-platform key; MINIMAX_TTS_MODEL / _SPEED /
  _VOL / _PITCH checked up front and recorded for --resume; --instruct and
  [direction] count only as one of MiniMax's emotion words.
- subtitle_type word → script-spelled words; --join block|all places the
  lines by them, so a joined minimax run needs no Gemini. The words are
  kept beside the audio for --resume.
- Pause marks <#0.4#> are voiced by minimax and dropped everywhere else;
  @pronounce lines become pronunciation_dict.tone.
- Docs: tools/audio/README (provider table, a MiniMax section with the
  year pitfall), playbook/04, bin/vh help and doctor, type 02, CHANGELOG.
- Other providers unchanged: say per line (zh, en) and with --join block
  give byte-identical timelines and audio with the old and new tts.py.
- A pause mark at a line's edge now adds to the pause between it and
  the next line between requests too (per line, --join block), not only
  inside one joined request; the README's own example relied on it.
- Spoken |x| outside a dialogue stays in the words (the subtitles were
  matched against text with reactions stripped).
- The 10,000-character limit is checked before the last take is deleted.
- --gap is recorded for a joined minimax run; --resume refuses a change.
- One spacing rule for dropped and merged pause marks: no space appears
  between a CJK character and a Latin word, and a mark under 0.01 s
  leaves the same text in the request as in the caption.
- "@pronounce" without a "/" stays a narration line with that id.
- Docs: volume range, what --resume keeps, the joined-run failure note,
  CHANGELOG wording.
@ZLHad
ZLHad marked this pull request as ready for review October 7, 2026 12:39
@ZLHad
ZLHad enabled auto-merge (squash) October 7, 2026 12:39
@ZLHad
ZLHad merged commit 56235f5 into main Oct 7, 2026
2 checks passed
@ZLHad
ZLHad deleted the claude/minimax-tts branch October 7, 2026 12:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant