Repository navigation
bin/vh tts: a minimax provider (MiniMax T2A v2), timed by its own subtitles - #69
Merged
Merged
Conversation
…titles A film's narration was made with MiniMax speech-2.8-hd through a project script: the whole script in one request kept the voice from drifting between blocks, and the request's own subtitles gave every character's time without an ASR pass. This makes it a provider. - POST https://api.minimax.io/v1/t2a_v2; key from MINIMAX_TOKEN_PLAN_KEY, then MINIMAX_API_KEY (never printed); hints for 1008 and 1004/2049; MINIMAX_BASE_URL for a China-platform key; MINIMAX_TTS_MODEL / _SPEED / _VOL / _PITCH checked up front and recorded for --resume; --instruct and [direction] count only as one of MiniMax's emotion words. - subtitle_type word → script-spelled words; --join block|all places the lines by them, so a joined minimax run needs no Gemini. The words are kept beside the audio for --resume. - Pause marks <#0.4#> are voiced by minimax and dropped everywhere else; @pronounce lines become pronunciation_dict.tone. - Docs: tools/audio/README (provider table, a MiniMax section with the year pitfall), playbook/04, bin/vh help and doctor, type 02, CHANGELOG. - Other providers unchanged: say per line (zh, en) and with --join block give byte-identical timelines and audio with the old and new tts.py.
- A pause mark at a line's edge now adds to the pause between it and the next line between requests too (per line, --join block), not only inside one joined request; the README's own example relied on it. - Spoken |x| outside a dialogue stays in the words (the subtitles were matched against text with reactions stripped). - The 10,000-character limit is checked before the last take is deleted. - --gap is recorded for a joined minimax run; --resume refuses a change. - One spacing rule for dropped and merged pause marks: no space appears between a CJK character and a Latin word, and a mark under 0.01 s leaves the same text in the request as in the caption. - "@pronounce" without a "/" stays a narration line with that id. - Docs: volume range, what --resume keeps, the joined-run failure note, CHANGELOG wording.
ZLHad
marked this pull request as ready for review
October 7, 2026 12:39
ZLHad
enabled auto-merge (squash)
October 7, 2026 12:39
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
A film's narration (2026-10-07) was made with MiniMax T2A
speech-2.8-hdthrough a project script. Two things made it worth a provider:subtitle_type: word) gave every character's time, so no ASR pass was needed for line boundaries or word captions.What changed
tools/audio/tts.py: new providerminimaxPOST https://api.minimax.io/v1/t2a_v2(the China hostapi.minimaxi.comrejects international keys with 2049;MINIMAX_BASE_URLfor a China-platform key).MINIMAX_TOKEN_PLAN_KEY(Token Plan, spends its credits) first, thenMINIMAX_API_KEY(pay-as-you-go). Errors name the variable, never the key, with a hint for 1008 (no balance on that billing) and 1004/2049 (key and host don't match). Checked before anything is deleted.MINIMAX_TTS_MODEL(defaultspeech-2.8-hd),MINIMAX_TTS_SPEED/_VOL/_PITCH: range-checked up front and recorded invo/<lang>/_run.json, so--resumerefuses a change.text_normalizationon,language_boostfrom--lang, 44.1 kHz WAV converted to 48 kHz like every provider.--instruct/[direction]count only when one of MiniMax's nine emotion words; anything else is ignored, said once.@pronouncename its pinyin, "Claude" comes as C / la / ude) and found in order in the text sent, then cut into the same script-spelled units elevenlabs gives. The file's character offsets are not used: one response counted the pause marks in them, another did not.--join block|allwith minimax places the lines by those words: no--align gemini, noGEMINI_API_KEY.--align geministill works (one transcription per block for the similarity check; times stay MiniMax's). Words are kept beside the audio (NN.words.json,_blockNN.words.json);--resumetreats a line or block as done only with them.<#0.4#>: voiced by minimax; captions and every other provider (gemini included) drop them. In a joined request lines are linked by<#--gap#>plus marks written at line edges; marks at a request's edges are dropped and runs merged, as MiniMax requires.@pronounce 硖合/(xia2)(he2)lines inscript.txt→pronunciation_dict.tone(minimax only; other providers print a note). A changed dictionary makes--resumerefuse.Docs:
tools/audio/README.md(provider table; a "MiniMax" section: keys and hosts, voices and how to list them viaPOST /v1/get_voice, settings, emotion words, word timing, one request for the whole script, pause marks,@pronounce, billing, measurements, and the year pitfall: "1953 年" is read 一千九百五十三年, so write 一九五三年; 2001/2003 read 两千零一/两千零三, acceptable);playbook/04-audio.md(commands, provider table, who can act, tags, 写念法, selection table);bin/vhhelp and doctor; type 02;captions.pydocstring;CHANGELOG.md.Known limitation (documented): MiniMax's Chinese voice ids contain a space, so they can't go into
@speakers; a minimax dialogue isn't possible yet. The default stays--join nonelike every other provider; the docs' examples use--join all.How it was verified
Against the live API (Token Plan key, 2026-10-07):
--join all(3 zh lines, 1 request, ~9 s): line spans and per-characterwordsfrom the subtitles,@pronouncesent, an inline<#0.4#>gave ~0.8 s with the comma; no Gemini call.calm; 平静地讲(emotion used, the rest noted once),--join blockwith a non-default voice andMINIMAX_TTS_SPEED=1.2, English side withhappy(3.5andGHz.timed separately).--resume: a finished run re-synthesizes nothing; a deleted block, or a block whose words file is missing, is synthesized again (the other block's audio unchanged, same md5); a changed speed or@pronounceentry is refused, naming what changed.MINIMAX_BASE_URL=https://api.minimaxi.com→ 2049 with the host hint; no key and a speed out of range exit before anything is deleted.--join all --align gemini: Gemini's daily transcription quota was used up, which exercised the failure path (lines still cut and timed by MiniMax,asr: {error}, take written, exit 1, no "block kept whole" message). The success path was run from a cached.asr.jsonwith--resume: per-line similarity and flag, head/tail at the block edges, times unchanged.bin/vh captionson the result:captions.jsonitems carrywords,.lines.srtwritten.Other providers unchanged (old
tts.pyfrommainvs new, same script with tags, reactions, bilingual lines):sayper line, zh and en: timeline, voiceover and_run.jsonbyte-identical (one first en run ofsaydiffered from every later run, old or new; repeated old runs match the new one);say --join blockfrom the same cached transcription: timeline, voiceover and every cut file byte-identical; with transcription failing (429): same output and timeline.tools/ci.sh --committedpasses, also withVH_BASH=/bin/bash(shellcheck and pyflakes through the uv shim).Independent review
One reviewer (fresh context) re-ran the old-vs-new comparisons with stubs and found no blocking bug. Fixed in the second commit, each re-tested against the API or offline:
--join block), though the README's example relied on it: it now adds to the silence there too (per line: exactly--gap+ mark between the two lines);|x|was missing fromwords;--gapis part of a joined minimax request: recorded in_run.json,--resumerefuses a change;@pronouncewithout a/stays a narration line with that id, as before;Not changed: a near-limit
--join allreturns its WAV as hex in the JSON (about 80 MB for a 457 s film, fine at our lengths).After the fixes,
sayper line (zh, en,--beats --snap downbeat,--gap/--lead) and--join blockwith--beatsare again byte-identical old vs new;tools/ci.sh --committedpasses (alsoVH_BASH=/bin/bash).