Merge pull request #562 from superdesigndev/codex/minimax-tts

feat(catalog): add MiniMax voice generation
This commit is contained in:
Taus
2026-09-18 04:37:30 +06:00
committed by GitHub
25 changed files with 371 additions and 60 deletions
File diff suppressed because one or more lines are too long
+23 -6
View File
@@ -137,6 +137,7 @@ sources:
- src/treg/catalog/examples/minimax.video-gen.result.retrieve.json
- src/treg/catalog/examples/minimax.video-gen.from_image.json
- src/treg/catalog/examples/minimax.video-gen.task.status.json
- src/treg/catalog/examples/minimax.voice-gen.voices.list.json
- src/treg/catalog/openrouter.yaml
- src/treg/catalog/openrouter.extended.yaml
- src/treg/catalog/examples/openrouter.x.alibaba-wan-3-0.json
@@ -575,8 +576,8 @@ The core AIGC generation rows pin `domain: models` too and carry PER-MODEL capab
(`video-gen.hailuo.from_text`, proposed in their provider files) rather than the job-level
`video-gen.from_text` family. Generation models are not interchangeable - a merged row comparing
Hailuo with Wan or Seedance is a false comparison - so the job-level capabilities are deliberately
memberless, reserved for hand-picked models (see capabilities.yaml). Both AI generation pages
therefore render as ONE flat model wall; the same model reachable over several routes (MiniMax
memberless, reserved for hand-picked models (see capabilities.yaml). The AI generation modality
pages therefore render as flat model walls; the same model reachable over several routes (MiniMax
direct, OpenRouter, Replicate all serve Hailuo) sits adjacent under model-led names, which is the
comparison that actually means something. The per-model capability is the join key that lets those
routes merge onto one row if that comparison is later curated. reAPI and PiAPI are the first pair
@@ -607,6 +608,7 @@ platforms:
web: "The web at large (backlinks, authority, traffic)"
video-gen: {label: "Video generation", category: "AI generation"}
image-gen: {label: "Image generation", category: "AI generation"}
voice-gen: {label: "Voice generation", category: "AI generation"}
```
Rules:
@@ -617,8 +619,8 @@ Rules:
this file. The validator accepts a capability that is either global or proposed in the same file.
- Under `AI generation`, platform means the generated-media modality rather than a system that owns
the data. The frozen vocabulary is `video-gen.from_text`, `video-gen.from_image`,
`video-gen.task.status`, `image-gen.from_text`, and `image-gen.edit`; text-to-video and
image-to-video stay separate because their required inputs and prices differ.
`video-gen.task.status`, `image-gen.from_text`, `image-gen.edit`, and `voice-gen.from_text`;
text-to-video and image-to-video stay separate because their required inputs and prices differ.
### `<service>.yaml`
@@ -673,7 +675,7 @@ endpoints:
### Async descriptors
`catalog_store._normalize` sets `cache: forbidden` for `image-gen` and `video-gen` endpoints,
`catalog_store._normalize` sets `cache: forbidden` for `image-gen`, `video-gen`, and `voice-gen` endpoints,
including synchronous generation, task/result utilities and generated extended rows. Their `kind`
is unchanged. These requests must reach the provider, not replay shared-account task ids or media
from an identical prompt. Other platforms retain their declared/default cache policy.
@@ -783,6 +785,17 @@ the terminal values `Success`/`Fail`, then pass the returned `file_id` to
`GET /v1/files/retrieve`. The v2 generation path serves the H3 family and is not a protocol upgrade
for the Hailuo models in this listing.
MiniMax also supplies the first `voice-gen` rows through the same provider connection. Speech 2.8
HD and Turbo are separate model rows over `POST /v1/t2a_v2`, each fixing its model plus
`stream: false` and `output_format: url`; this keeps the response bounded and returns a 24-hour
audio URL. The route is synchronous and billed per input character, so `_body_text_characters`
scales the reserve from the provider-facing `text` value. The two rows remain marked `skipped`
until a deliberate paid verification call is authorized; documentation provenance is enough for
platform eligibility, but is not presented as live route evidence. The Voice generation Actions
shelf exposes `minimax.voice-gen.voices.list` so callers can discover valid system voice IDs. Its
request is fixed to `voice_type: system`: account-specific cloned and generated voices must not
cross team boundaries when treg's shared MiniMax connection is used.
reAPI answers every submission with a bare `{id, status}` and reports the charge on the poll body
(`usage.credits`, 1 credit = $0.001); video rows keep the file-level descriptor (`output.video_urls`)
and image rows replace it whole for `output.image_urls`. PiAPI wraps its task routes in
@@ -1046,7 +1059,11 @@ that slice: it bills one whole credit (~$0.0245) for one email or ten (observed
10 emails"); `Catalog.advertised_usd` is what `catalog_search` / `catalog_get` put on
`usd_per_call`. Settlement still reads `usd` and the derived email-count rule — display only.
Akta bills 1.5 credits per 50 reviews the same `per` way. Without `per`, every one of those had
to be either wrong or rounded into prose.
to be either wrong or rounded into prose. A scalar `unit: character` is request-priced rather than
page-priced: `_body_text_characters` counts the top-level JSON `text` string and multiplies the
normalized per-character USD rate. Invalid JSON or a missing/empty string reserves one character,
never zero; the normal request/envelope checks decide whether the provider served anything and
`per_success` releases a rejected call.
**Three kinds of denomination convert, and they convert differently:**
+5 -2
View File
@@ -648,10 +648,13 @@ platform without opening it, and every field comes off the `/catalog/platforms`
provider with no published rate stays silent. Note that `price_from` arrives as `null` *or* as an
empty `{}`, and the empty object has to be normalised to null first — being truthy, it otherwise
short-circuits the auth-kind branch and silently costs an OAuth-only platform its "free with your
account".
account". A grouped scalar rate such as MiniMax TTS supplies `display_usd` plus `display_unit`, so
the card says `$0.60 / 10000 characters` instead of rounding the normalized per-character rate
down to an unreadable number.
**Prices are unified USD.** Every price the marketplace displays — the card footer, the capability card's
"from", and the per-endpoint cost chip — is the **server's computed `usd`** field on `cost` / `price_from`,
"from", and the per-endpoint cost chip — comes from the server's price object on `cost` / `price_from`:
normally its computed **`usd`**, or its equivalent grouped **`display_usd` / `display_unit`** pair,
formatted by `usdNum`: two significant figures under a dollar (`$0.015`, `$0.00015`), cents at or above one.
The FX table lives in the catalog (`fx.yaml`) so a rate refresh re-prices every surface at once, and the
dashboard carries **no** conversion constant of its own — one here would drift from the CLI the moment the
+6 -4
View File
@@ -36,10 +36,12 @@ One skill, three personas:
- **consumer** — discover + call tools with no credentials locally. Teaches the agent-native
**URL-passthrough** first: take the real upstream URL and prefix it with `{BASE}/call/`
+ the `X-Treg-Token` header; `treg call <tool> <path>` is the CLI shorthand. Its
"generate a video or an image" task teaches the async shape: `--await` for the CLI, the
`X-Treg-Async` descriptor and lazy 30-60 s polling for MCP agents, the shell-timeout warning
(video takes 1-5 minutes), reserve-then-settle money with refunds on failure, and expiring
result URLs that treg never stores.
generation task distinguishes synchronous MiniMax voice generation from asynchronous video and
image generation. Voice callers discover current system voice IDs through
`minimax.voice-gen.voices.list`, then call the HD or Turbo endpoint for a temporary audio URL.
Video/image callers learn `--await`, the `X-Treg-Async` descriptor and lazy 30-60 s polling for
MCP agents, the shell-timeout warning (video takes 1-5 minutes), reserve-then-settle money with
releases on failure, and expiring result URLs that treg never stores.
- **creator** — turn a local skill into a shared tool: `treg secret add`, `treg tool add` (single-key or
`--bind` multi-credential), the `treg skill scaffold → push` bundle flow, and `treg oauth connect` for
browser-consent tokens. Documents the two OAuth modes (auto-refresh vs manual) and the four auth shapes.
+13 -5
View File
@@ -188,11 +188,11 @@ Notes:
verify call each would have caught.
- An endpoint with no published price is refused rather than served free; connect your own key.
## Task - generate a video or an image
## Task - generate video, images, or voice
Generation models live in the catalog under the `video-gen` and `image-gen` platforms, one row per
model per route (MiniMax direct, Replicate, OpenRouter), so the same model on two routes sits next
to itself with both prices. Models are not interchangeable - you pick one; treg does not choose.
Generation models live in the catalog under the `video-gen`, `image-gen`, and `voice-gen` platforms,
one row per model per route, so the same model on two routes sits next to itself with both prices.
Models are not interchangeable - you pick one; treg does not choose.
```bash
treg catalog search "text to video" # every model, with prices
@@ -200,9 +200,17 @@ treg catalog get minimax.video-gen.h3.generate # native params, model enum
treg call minimax.video-gen.h3.generate --await --timeout 900 --data '{"model":"MiniMax-H3-Max",
"content":[{"type":"text","text":"A paper boat drifts across a quiet pond at sunrise."}],
"resolution":"480P","duration":5,"ratio":"16:9"}'
treg call minimax.voice-gen.voices.list --data '{"voice_type":"system"}'
treg catalog get minimax.voice-gen.speech-2-8-turbo
treg call minimax.voice-gen.speech-2-8-turbo --data '{"model":"speech-2.8-turbo",
"text":"A calm voice can make a complex idea feel simple.","stream":false,"output_format":"url",
"voice_setting":{"voice_id":"English_expressive_narrator","speed":1,"vol":1,"pitch":0}}'
```
How it works:
- **A generation call is an async task.** The submission returns a task id at once; `--await` polls
- **Voice generation is synchronous.** MiniMax returns JSON containing a 24-hour audio URL. The
catalog route fixes `stream:false` and `output_format:"url"`; use the voice-list action to discover
valid system voice IDs, then choose HD or Turbo by endpoint id.
- **A video or image generation call is an async task.** The submission returns a task id at once; `--await` polls
the provider until it finishes and prints the **final response only** on stdout. stderr carries the
task id, a resumable `treg call …` command (Ctrl-C loses the wait, never the task or the money),
progress, and the result URL. Exit 0 = done, 2 = the provider failed the task, 3 = timed out
+13 -5
View File
@@ -191,11 +191,11 @@ Notes:
verify call each would have caught.
- An endpoint with no published price is refused rather than served free; connect your own key.
## Task - generate a video or an image
## Task - generate video, images, or voice
Generation models live in the catalog under the `video-gen` and `image-gen` platforms, one row per
model per route (MiniMax direct, Replicate, OpenRouter), so the same model on two routes sits next
to itself with both prices. Models are not interchangeable - you pick one; treg does not choose.
Generation models live in the catalog under the `video-gen`, `image-gen`, and `voice-gen` platforms,
one row per model per route, so the same model on two routes sits next to itself with both prices.
Models are not interchangeable - you pick one; treg does not choose.
```bash
treg catalog search "text to video" # every model, with prices
@@ -203,9 +203,17 @@ treg catalog get minimax.video-gen.h3.generate # native params, model enum
treg call minimax.video-gen.h3.generate --await --timeout 900 --data '{"model":"MiniMax-H3-Max",
"content":[{"type":"text","text":"A paper boat drifts across a quiet pond at sunrise."}],
"resolution":"480P","duration":5,"ratio":"16:9"}'
treg call minimax.voice-gen.voices.list --data '{"voice_type":"system"}'
treg catalog get minimax.voice-gen.speech-2-8-turbo
treg call minimax.voice-gen.speech-2-8-turbo --data '{"model":"speech-2.8-turbo",
"text":"A calm voice can make a complex idea feel simple.","stream":false,"output_format":"url",
"voice_setting":{"voice_id":"English_expressive_narrator","speed":1,"vol":1,"pitch":0}}'
```
How it works:
- **A generation call is an async task.** The submission returns a task id at once; `--await` polls
- **Voice generation is synchronous.** MiniMax returns JSON containing a 24-hour audio URL. The
catalog route fixes `stream:false` and `output_format:"url"`; use the voice-list action to discover
valid system voice IDs, then choose HD or Turbo by endpoint id.
- **A video or image generation call is an async task.** The submission returns a task id at once; `--await` polls
the provider until it finishes and prints the **final response only** on stdout. stderr carries the
task id, a resumable `treg call …` command (Ctrl-C loses the wait, never the task or the money),
progress, and the result URL. Exit 0 = done, 2 = the provider failed the task, 3 = timed out
+13 -5
View File
@@ -172,11 +172,11 @@ Notes:
verify call each would have caught.
- An endpoint with no published price is refused rather than served free; connect your own key.
## Task - generate a video or an image
## Task - generate video, images, or voice
Generation models live in the catalog under the `video-gen` and `image-gen` platforms, one row per
model per route (MiniMax direct, Replicate, OpenRouter), so the same model on two routes sits next
to itself with both prices. Models are not interchangeable - you pick one; treg does not choose.
Generation models live in the catalog under the `video-gen`, `image-gen`, and `voice-gen` platforms,
one row per model per route, so the same model on two routes sits next to itself with both prices.
Models are not interchangeable - you pick one; treg does not choose.
```bash
treg catalog search "text to video" # every model, with prices
@@ -184,9 +184,17 @@ treg catalog get minimax.video-gen.h3.generate # native params, model enum
treg call minimax.video-gen.h3.generate --await --timeout 900 --data '{"model":"MiniMax-H3-Max",
"content":[{"type":"text","text":"A paper boat drifts across a quiet pond at sunrise."}],
"resolution":"480P","duration":5,"ratio":"16:9"}'
treg call minimax.voice-gen.voices.list --data '{"voice_type":"system"}'
treg catalog get minimax.voice-gen.speech-2-8-turbo
treg call minimax.voice-gen.speech-2-8-turbo --data '{"model":"speech-2.8-turbo",
"text":"A calm voice can make a complex idea feel simple.","stream":false,"output_format":"url",
"voice_setting":{"voice_id":"English_expressive_narrator","speed":1,"vol":1,"pitch":0}}'
```
How it works:
- **A generation call is an async task.** The submission returns a task id at once; `--await` polls
- **Voice generation is synchronous.** MiniMax returns JSON containing a 24-hour audio URL. The
catalog route fixes `stream:false` and `output_format:"url"`; use the voice-list action to discover
valid system voice IDs, then choose HD or Turbo by endpoint id.
- **A video or image generation call is an async task.** The submission returns a task id at once; `--await` polls
the provider until it finishes and prints the **final response only** on stdout. stderr carries the
task id, a resumable `treg call …` command (Ctrl-C loses the wait, never the task or the money),
progress, and the result URL. Exit 0 = done, 2 = the provider failed the task, 3 = timed out
+13 -5
View File
@@ -186,11 +186,11 @@ Notes:
verify call each would have caught.
- An endpoint with no published price is refused rather than served free; connect your own key.
## Task - generate a video or an image
## Task - generate video, images, or voice
Generation models live in the catalog under the `video-gen` and `image-gen` platforms, one row per
model per route (MiniMax direct, Replicate, OpenRouter), so the same model on two routes sits next
to itself with both prices. Models are not interchangeable - you pick one; treg does not choose.
Generation models live in the catalog under the `video-gen`, `image-gen`, and `voice-gen` platforms,
one row per model per route, so the same model on two routes sits next to itself with both prices.
Models are not interchangeable - you pick one; treg does not choose.
```bash
treg catalog search "text to video" # every model, with prices
@@ -198,9 +198,17 @@ treg catalog get minimax.video-gen.h3.generate # native params, model enum
treg call minimax.video-gen.h3.generate --await --timeout 900 --data '{"model":"MiniMax-H3-Max",
"content":[{"type":"text","text":"A paper boat drifts across a quiet pond at sunrise."}],
"resolution":"480P","duration":5,"ratio":"16:9"}'
treg call minimax.voice-gen.voices.list --data '{"voice_type":"system"}'
treg catalog get minimax.voice-gen.speech-2-8-turbo
treg call minimax.voice-gen.speech-2-8-turbo --data '{"model":"speech-2.8-turbo",
"text":"A calm voice can make a complex idea feel simple.","stream":false,"output_format":"url",
"voice_setting":{"voice_id":"English_expressive_narrator","speed":1,"vol":1,"pitch":0}}'
```
How it works:
- **A generation call is an async task.** The submission returns a task id at once; `--await` polls
- **Voice generation is synchronous.** MiniMax returns JSON containing a 24-hour audio URL. The
catalog route fixes `stream:false` and `output_format:"url"`; use the voice-list action to discover
valid system voice IDs, then choose HD or Turbo by endpoint id.
- **A video or image generation call is an async task.** The submission returns a task id at once; `--await` polls
the provider until it finishes and prints the **final response only** on stdout. stderr carries the
task id, a resumable `treg call …` command (Ctrl-C loses the wait, never the task or the money),
progress, and the result URL. Exit 0 = done, 2 = the provider failed the task, 3 = timed out
+13 -5
View File
@@ -170,11 +170,11 @@ Notes:
verify call each would have caught.
- An endpoint with no published price is refused rather than served free; connect your own key.
## Task - generate a video or an image
## Task - generate video, images, or voice
Generation models live in the catalog under the `video-gen` and `image-gen` platforms, one row per
model per route (MiniMax direct, Replicate, OpenRouter), so the same model on two routes sits next
to itself with both prices. Models are not interchangeable - you pick one; treg does not choose.
Generation models live in the catalog under the `video-gen`, `image-gen`, and `voice-gen` platforms,
one row per model per route, so the same model on two routes sits next to itself with both prices.
Models are not interchangeable - you pick one; treg does not choose.
```bash
treg catalog search "text to video" # every model, with prices
@@ -182,9 +182,17 @@ treg catalog get minimax.video-gen.h3.generate # native params, model enum
treg call minimax.video-gen.h3.generate --await --timeout 900 --data '{"model":"MiniMax-H3-Max",
"content":[{"type":"text","text":"A paper boat drifts across a quiet pond at sunrise."}],
"resolution":"480P","duration":5,"ratio":"16:9"}'
treg call minimax.voice-gen.voices.list --data '{"voice_type":"system"}'
treg catalog get minimax.voice-gen.speech-2-8-turbo
treg call minimax.voice-gen.speech-2-8-turbo --data '{"model":"speech-2.8-turbo",
"text":"A calm voice can make a complex idea feel simple.","stream":false,"output_format":"url",
"voice_setting":{"voice_id":"English_expressive_narrator","speed":1,"vol":1,"pitch":0}}'
```
How it works:
- **A generation call is an async task.** The submission returns a task id at once; `--await` polls
- **Voice generation is synchronous.** MiniMax returns JSON containing a 24-hour audio URL. The
catalog route fixes `stream:false` and `output_format:"url"`; use the voice-list action to discover
valid system voice IDs, then choose HD or Turbo by endpoint id.
- **A video or image generation call is an async task.** The submission returns a task id at once; `--await` polls
the provider until it finishes and prints the **final response only** on stdout. stderr carries the
task id, a resumable `treg call …` command (Ctrl-C loses the wait, never the task or the money),
progress, and the result URL. Exit 0 = done, 2 = the provider failed the task, 3 = timed out
+20 -1
View File
@@ -465,6 +465,23 @@ def _body_limit(body: bytes) -> int | None:
return items
def _body_text_characters(body: bytes) -> int:
"""Count the provider-facing ``text`` field for character-priced generation calls.
The catalog price is already normalized to USD per character. Invalid JSON or a missing text
field reserves one unit rather than zero; platform request validation/provider rejection still
decides whether the call is relayed or charged.
"""
if body:
try:
document = json.loads(body)
except (ValueError, UnicodeDecodeError):
document = None
if isinstance(document, dict) and isinstance(document.get("text"), str):
return max(1, len(document["text"]))
return 1
def _platform_estimate_micro(cost: dict, query, body: bytes = b"") -> int:
"""What one call is expected to cost the platform, in RAW micro-USD (no margin — ledger.reserve
applies that). Rounds UP: a fraction of a micro-dollar is not representable and must not round to
@@ -473,7 +490,9 @@ def _platform_estimate_micro(cost: dict, query, body: bytes = b"") -> int:
if usd is None:
return 0
n = 1
if cost.get("type") in ("per_result", "quota_rows") and cost.get("unit") in _ENTITY_UNITS:
if cost.get("unit") == "character":
n = _body_text_characters(body)
elif cost.get("type") in ("per_result", "quota_rows") and cost.get("unit") in _ENTITY_UNITS:
# Priced per INPUT entity, not per returned row: the page-size default below has no
# meaning here and billed one-target calls 20x (seranking summary, serpstat overview —
# 2026-09-05). The request names how many entities it asks about.
+2
View File
@@ -30,6 +30,8 @@ aliases:
i2v: [image-to-video]
texttovideo: [text-to-video]
imagetovideo: [image-to-video]
tts: [text-to-speech]
texttospeech: [text-to-speech]
# sales vocabulary — the catalog says "people" / "contact"; agents say "leads" / "prospects"
# (logged miss 2026-08-28: `find leads` matched only the provider names leadsforge/predictleads)
lead: [leads]
+2
View File
@@ -28,6 +28,7 @@ platforms:
# --- AI generation: generated media modalities ------------------------------------------------
video-gen: {label: "Video generation", category: "AI generation", summary: "Text-to-video and image-to-video across models, with prices side by side."}
image-gen: {label: "Image generation", category: "AI generation", summary: "Text-to-image and prompt-based image editing across models."}
voice-gen: {label: "Voice generation", category: "AI generation", summary: "Text-to-speech across voice models, with prices side by side."}
# --- Social: global social & video -----------------------------------------------------------
tiktok: {label: "TikTok", category: "Social", featured: 1, summary: "Profiles, videos, comments, hashtags and search — plus your own account via OAuth."}
@@ -135,6 +136,7 @@ capabilities:
video-gen.task.status: "Check an async video generation task"
image-gen.from_text: "Generate an image from a text prompt"
image-gen.edit: "Edit or restyle an image with a prompt"
voice-gen.from_text: "Generate speech from text"
creators.search: "Find influencers & creators by niche, country, followers & engagement on Instagram, TikTok, YouTube"
# --- SEO / search -----------------------------------------------------------------------------
@@ -0,0 +1,24 @@
{
"system_voice": [
{
"voice_id": "English_expressive_narrator",
"description": [
"An expressive adult male voice with a British accent, perfect for engaging audiobook narration."
],
"voice_name": "Expressive Narrator",
"created_time": "2025-01-01"
},
{
"voice_id": "English_radiant_girl",
"description": [
"A radiant and lively young adult female voice with a general American accent, full of energy and brightness."
],
"voice_name": "Radiant Girl",
"created_time": "2025-01-01"
}
],
"base_resp": {
"status_code": 0,
"status_msg": "success"
}
}
+131 -1
View File
@@ -3,7 +3,7 @@ source:
docs: https://platform.minimax.io/docs/api-reference/api-overview
curated: 2026-09-02
pricing_url: https://platform.minimax.io/docs/guides/pricing-paygo
limits: "Video generation is asynchronous; poll every 10 seconds. Generated download URLs expire after about 9 hours."
limits: "Video generation is asynchronous; poll every 10 seconds. Video download URLs expire after about 9 hours. Synchronous TTS accepts up to 10,000 characters; generated audio URLs expire after 24 hours."
proposed_capabilities:
video-gen.result.retrieve: "Retrieve completed video generation content"
# Per-model capabilities: each is the join key for one model, so the row stands alone in the
@@ -14,6 +14,9 @@ proposed_capabilities:
video-gen.hailuo.from_image: "Animate an image into a Hailuo video"
video-gen.h3.generate: "Generate an H3 (Hailuo 3) video from text, optionally with first/last frame images"
image-gen.image-01.from_text: "Generate images with MiniMax image-01"
voice-gen.speech-2-8-hd.generate: "MiniMax Speech 2.8 HD"
voice-gen.speech-2-8-turbo.generate: "MiniMax Speech 2.8 Turbo"
voice-gen.voices.list: "List available MiniMax voices"
async:
id_from: task_id
poll:
@@ -284,3 +287,130 @@ endpoints:
confidence: documented
note: "Price is per generated image."
docs_url: https://platform.minimax.io/docs/api-reference/image-generation-t2i
- id: minimax.voice-gen.speech-2-8-hd
capability: voice-gen.speech-2-8-hd.generate
platform: voice-gen
domain: models
scope: any_account
method: POST
path: /v1/t2a_v2
name: "MiniMax: Speech 2.8 HD"
summary: "Generate high-definition speech from up to 10,000 characters of text."
async: false
platform_request:
body.model: speech-2.8-hd
body.stream: false
body.output_format: url
input:
body:
model: {type: string, required: true, enum: [speech-2.8-hd], example: speech-2.8-hd}
text: {type: string, required: true, max: 10000, example: "A calm voice can make a complex idea feel simple.", note: "Fewer than 10,000 characters. Use <#x#> between speakable segments for a pause of x seconds."}
stream: {type: boolean, required: true, enum: [false], example: false, note: "This catalog route is non-streaming so the relay returns one bounded JSON response."}
output_format: {type: string, required: true, enum: [url], example: url, note: "The returned audio URL is valid for 24 hours."}
voice_setting: {type: object, required: true, example: {voice_id: English_expressive_narrator, speed: 1, vol: 1, pitch: 0}, note: "voice_id is required. Optional controls include speed, vol, pitch, emotion, text_normalization and latex_read; see the API reference for ranges."}
audio_setting: {type: object, required: false, example: {sample_rate: 32000, bitrate: 128000, format: mp3, channel: 1}, note: "Non-streaming output supports mp3, wav and flac."}
language_boost: {type: string, required: false, example: auto, note: "Set auto for automatic language detection, or use one of MiniMax's documented language names."}
pronunciation_dict: {type: object, required: false, example: {tone: ["Treg/tree-g"]}, note: "Custom pronunciation pairs in source/target form."}
subtitle_enable: {type: boolean, required: false, default: false}
subtitle_type: {type: string, required: false, default: sentence, enum: [sentence, word], example: sentence}
voice_modify: {type: object, required: false, note: "Optional pitch, intensity, timbre and sound_effects controls."}
bodyType: json
test_request:
body:
model: speech-2.8-hd
text: "Hello."
stream: false
output_format: url
voice_setting: {voice_id: English_expressive_narrator, speed: 1, vol: 1, pitch: 0}
audio_setting: {sample_rate: 32000, bitrate: 128000, format: mp3, channel: 1}
expect: {json_path: base_resp.status_code, equals: 0}
cost:
type: per_success
value: 1.00
currency: USD
per: 10000
unit: character
display: {unit: characters, grouped: true}
source: docs
source_url: https://platform.minimax.io/docs/guides/pricing-paygo#audio
checked: 2026-09-18
confidence: documented
note: "MiniMax bills synchronous Speech 2.8 HD by input characters at $100 per million ($1 per 10,000). Failed requests are not charged."
skipped: "Not live-called: synthesis spends real MiniMax credits, and no explicit paid verification was authorized for this change."
docs_url: https://platform.minimax.io/docs/api-reference/speech-t2a-http
- id: minimax.voice-gen.speech-2-8-turbo
capability: voice-gen.speech-2-8-turbo.generate
platform: voice-gen
domain: models
scope: any_account
method: POST
path: /v1/t2a_v2
name: "MiniMax: Speech 2.8 Turbo"
summary: "Generate fast, cost-efficient speech from up to 10,000 characters of text."
async: false
platform_request:
body.model: speech-2.8-turbo
body.stream: false
body.output_format: url
input:
body:
model: {type: string, required: true, enum: [speech-2.8-turbo], example: speech-2.8-turbo}
text: {type: string, required: true, max: 10000, example: "A calm voice can make a complex idea feel simple.", note: "Fewer than 10,000 characters. Use <#x#> between speakable segments for a pause of x seconds."}
stream: {type: boolean, required: true, enum: [false], example: false, note: "This catalog route is non-streaming so the relay returns one bounded JSON response."}
output_format: {type: string, required: true, enum: [url], example: url, note: "The returned audio URL is valid for 24 hours."}
voice_setting: {type: object, required: true, example: {voice_id: English_expressive_narrator, speed: 1, vol: 1, pitch: 0}, note: "voice_id is required. Optional controls include speed, vol, pitch, emotion, text_normalization and latex_read; see the API reference for ranges."}
audio_setting: {type: object, required: false, example: {sample_rate: 32000, bitrate: 128000, format: mp3, channel: 1}, note: "Non-streaming output supports mp3, wav and flac."}
language_boost: {type: string, required: false, example: auto, note: "Set auto for automatic language detection, or use one of MiniMax's documented language names."}
pronunciation_dict: {type: object, required: false, example: {tone: ["Treg/tree-g"]}, note: "Custom pronunciation pairs in source/target form."}
subtitle_enable: {type: boolean, required: false, default: false}
subtitle_type: {type: string, required: false, default: sentence, enum: [sentence, word], example: sentence}
voice_modify: {type: object, required: false, note: "Optional pitch, intensity, timbre and sound_effects controls."}
bodyType: json
test_request:
body:
model: speech-2.8-turbo
text: "Hello."
stream: false
output_format: url
voice_setting: {voice_id: English_expressive_narrator, speed: 1, vol: 1, pitch: 0}
audio_setting: {sample_rate: 32000, bitrate: 128000, format: mp3, channel: 1}
expect: {json_path: base_resp.status_code, equals: 0}
cost:
type: per_success
value: 0.60
currency: USD
per: 10000
unit: character
display: {unit: characters, grouped: true}
source: docs
source_url: https://platform.minimax.io/docs/guides/pricing-paygo#audio
checked: 2026-09-18
confidence: documented
note: "MiniMax bills synchronous Speech 2.8 Turbo by input characters at $60 per million ($0.60 per 10,000). Failed requests are not charged."
skipped: "Not live-called: synthesis spends real MiniMax credits, and no explicit paid verification was authorized for this change."
docs_url: https://platform.minimax.io/docs/api-reference/speech-t2a-http
- id: minimax.voice-gen.voices.list
kind: utility
capability: voice-gen.voices.list
platform: voice-gen
scope: any_account
method: POST
path: /v1/get_voice
name: "List available voices"
summary: "List the current MiniMax system voices and their voice IDs."
async: false
platform_request:
body.voice_type: system
input:
body:
voice_type: {type: string, required: true, enum: [system], example: system, note: "This catalog action intentionally exposes system voices only, so account-specific cloned or generated voices never cross team boundaries on treg's shared MiniMax connection."}
bodyType: json
test_request: {body: {voice_type: system}}
expect: {json_path: base_resp.status_code, equals: 0}
cost: {type: free, value: 0, currency: USD, unit: call, note: "Voice discovery is a free account-management read."}
verified: 2026-09-18
example_response: examples/minimax.voice-gen.voices.list.json
docs_url: https://platform.minimax.io/docs/api-reference/voice-management-get
+1 -1
View File
@@ -230,7 +230,7 @@ class Settings(BaseSettings):
platform_key_aviato: str = "" # Bearer key; $10 auto-top-up buys 1,000 credits
platform_key_exa: str = "" # x-api-key; dollar-metered ($7/1k searches, $1/1k pages); settles from costDollars.total
platform_key_cloro: str = "" # Bearer key (sk_live_…); Hobby metered rate $0.0004/credit; settles from X-Credits-Charged
platform_key_minimax: str = "" # Bearer key for asynchronous Hailuo generation
platform_key_minimax: str = "" # Bearer key for MiniMax voice, image and video generation
platform_key_openrouter: str = "" # Bearer key for asynchronous routed generation
platform_key_replicate: str = "" # Bearer token for official asynchronous models
platform_key_reapi: str = "" # Bearer key; prepaid credits at $0.001, Seedance 2.5 + image models
+1 -1
View File
@@ -660,7 +660,7 @@ def _normalize(raw: dict, provider: str, directory: Path) -> dict:
# inherited from the file header unless the endpoint declares its own.
# Generated media and its task/result utilities must never replay shared-account ids.
# Enforce this for core and generated extended rows without changing their public kind.
"cache": "forbidden" if platform in {"image-gen", "video-gen"} else raw.get("cache"),
"cache": "forbidden" if platform in {"image-gen", "video-gen", "voice-gen"} else raw.get("cache"),
"verified": str(verified) if verified else None,
# {status, means} — a status the provider uses for "asked and answered: no result" (PDL
# 404s a person it has no record of). Only endpoints with evidenced miss semantics carry
+1 -1
View File
@@ -1673,7 +1673,7 @@ MINIMAX = OAuthProvider(
auth_uri="", token_uri="", scopes={},
client_id_setting="", client_secret_setting="",
category="AI generation",
summary="Generate images and create videos from text or source images.",
summary="Generate voice, images, and videos from text or source images.",
base_url="https://api.minimax.io",
docs_url="https://platform.minimax.io/docs/api-reference/api-overview",
probe_path="/v2/video_generation",
+5 -2
View File
@@ -4094,7 +4094,7 @@ createApp({
// catalog invents tomorrow still gets a shelf, sorted to the end.
const order=['Enrichment','SEO/AEO','Social','Advertising','E-commerce','Reviews & Apps','AI generation','Community'];
const hints={
'AI generation':'video and image models, the same model over several routes priced side by side',
'AI generation':'video, image and voice models, the same model over several routes priced side by side',
'SEO/AEO':'rankings, keywords and backlinks — what search engines know, and what the answer engines say',
'Social':'posts, profiles and comments, straight from the feeds',
'Enrichment':'people and company records, resolved from an email or a domain',
@@ -5772,7 +5772,10 @@ createApp({
platPrice(pl){
const raw=pl && pl.price_from;
const pf=(raw && Object.keys(raw).length) ? raw : null;
const paid=(pf && typeof pf.usd==='number') ? '$'+this.usdNum(pf.usd)+' / '+this.priceUnit(pf.type) : null;
const paid=(pf && typeof pf.usd==='number')
? '$'+this.usdNum(typeof pf.display_usd==='number' ? pf.display_usd : pf.usd)+' / '+
(pf.display_unit || this.priceUnit(pf.type))
: null;
if(pf && pf.type==='free') return {free:true, text:'free with your account'};
// An OAuth integration among the providers means the floor price is $0: the account you
// connect IS the licence. Metered providers may serve the same platform (that rate moves to
+13 -7
View File
@@ -251,13 +251,19 @@ unverified rows and guesses that one $0.006 verify call each would have caught.)
Call a specific provider's endpoint when you need a particular one.
<!--/routed-->
**Video and image generation - async calls.** The `video-gen` and `image-gen` platforms list one
row per model per route (MiniMax direct, Replicate, OpenRouter); the same model on two routes sits
side by side with both prices. Models are not interchangeable - you pick one; treg does not choose.
A generation call is an **async task**: the submission returns a task id, and the response carries
`X-Treg-Async`, a JSON descriptor saying where to poll (a catalog id plus the parameter that takes
the task id, or an allow-listed URL from the submission), which status values mean done or failed,
and where the result is (a path in the terminal response, or one more catalog call that fetches it).
**Video, image and voice generation.** The `video-gen`, `image-gen`, and `voice-gen` platforms list
one row per model per route; the same model on two routes sits side by side with both prices. Models
are not interchangeable - you pick one; treg does not choose. MiniMax voice generation is
synchronous: call `minimax.voice-gen.speech-2-8-hd` or
`minimax.voice-gen.speech-2-8-turbo` with `stream:false` and `output_format:"url"`; the response
contains a generated-audio URL valid for 24 hours and the price scales with input characters.
Call `minimax.voice-gen.voices.list` with `{"voice_type":"system"}` to discover valid voice IDs.
**Video and image generation use async calls.** A generation call is an **async task**: the
submission returns a task id, and the response carries `X-Treg-Async`, a JSON descriptor saying
where to poll (a catalog id plus the parameter that takes the task id, or an allow-listed URL from
the submission), which status values mean done or failed, and where the result is (a path in the
terminal response, or one more catalog call that fetches it).
treg catalog get minimax.video-gen.h3.generate # native params, model enum, price table, descriptor
treg call minimax.video-gen.h3.generate --await --timeout 900 --data '{...}'
+13 -5
View File
@@ -151,11 +151,11 @@ Notes:
<!--/routed-->
- An endpoint with no published price is refused rather than served free; connect your own key.
## Task - generate a video or an image
## Task - generate video, images, or voice
Generation models live in the catalog under the `video-gen` and `image-gen` platforms, one row per
model per route (MiniMax direct, Replicate, OpenRouter), so the same model on two routes sits next
to itself with both prices. Models are not interchangeable - you pick one; treg does not choose.
Generation models live in the catalog under the `video-gen`, `image-gen`, and `voice-gen` platforms,
one row per model per route, so the same model on two routes sits next to itself with both prices.
Models are not interchangeable - you pick one; treg does not choose.
```bash
treg catalog search "text to video" # every model, with prices
@@ -163,9 +163,17 @@ treg catalog get minimax.video-gen.h3.generate # native params, model enum
treg call minimax.video-gen.h3.generate --await --timeout 900 --data '{"model":"MiniMax-H3-Max",
"content":[{"type":"text","text":"A paper boat drifts across a quiet pond at sunrise."}],
"resolution":"480P","duration":5,"ratio":"16:9"}'
treg call minimax.voice-gen.voices.list --data '{"voice_type":"system"}'
treg catalog get minimax.voice-gen.speech-2-8-turbo
treg call minimax.voice-gen.speech-2-8-turbo --data '{"model":"speech-2.8-turbo",
"text":"A calm voice can make a complex idea feel simple.","stream":false,"output_format":"url",
"voice_setting":{"voice_id":"English_expressive_narrator","speed":1,"vol":1,"pitch":0}}'
```
How it works:
- **A generation call is an async task.** The submission returns a task id at once; `--await` polls
- **Voice generation is synchronous.** MiniMax returns JSON containing a 24-hour audio URL. The
catalog route fixes `stream:false` and `output_format:"url"`; use the voice-list action to discover
valid system voice IDs, then choose HD or Turbo by endpoint id.
- **A video or image generation call is an async task.** The submission returns a task id at once; `--await` polls
the provider until it finishes and prints the **final response only** on stdout. stderr carries the
task id, a resumable `treg call …` command (Ctrl-C loses the wait, never the task or the money),
progress, and the result URL. Exit 0 = done, 2 = the provider failed the task, 3 = timed out
+1 -1
View File
@@ -37,7 +37,7 @@ EP = "replicate.image-gen.flux-schnell"
def test_all_generation_catalog_entries_forbid_cache_including_extended():
entries = [ep for ep in catalog_store.load().endpoints
if ep["platform"] in {"image-gen", "video-gen"}]
if ep["platform"] in {"image-gen", "video-gen", "voice-gen"}]
assert any(".x." in ep["id"] for ep in entries)
assert any(ep["id"] == "minimax.image-gen.from_text" for ep in entries)
for ep in entries:
+38 -1
View File
@@ -891,8 +891,13 @@ def test_ai_generation_taxonomy_and_chinese_alias_tokens_are_loaded():
cat = cs.load()
assert cat.platforms["video-gen"]["category"] == "AI generation"
assert cat.platforms["image-gen"]["category"] == "AI generation"
assert cat.platforms["voice-gen"] == {
"label": "Voice generation",
"category": "AI generation",
"summary": "Text-to-speech across voice models, with prices side by side.",
}
assert {"video-gen.from_text", "video-gen.from_image", "video-gen.task.status",
"image-gen.from_text", "image-gen.edit"} <= set(cat.capabilities)
"image-gen.from_text", "image-gen.edit", "voice-gen.from_text"} <= set(cat.capabilities)
text_to_video_zh = "\u6587\u751f\u89c6\u9891"
assert cat.aliases[text_to_video_zh] == ["text-to-video"]
assert cs._tokens(f"{text_to_video_zh} text-to-video") == [
@@ -940,6 +945,38 @@ async def test_ai_generation_pages_keep_comparisons_curated_and_coverage_in_mode
assert {"minimax.image-gen.from_text", "replicate.image-gen.flux-schnell",
"reapi.image-gen.gemini-3-pro-image", "piapi.image-gen.gpt-image-2-5"} <= image_ids
voice = (await clients.get("/catalog/platforms/voice-gen")).json()
assert {section["domain"] for section in voice["domains"]} == {"models"}
voice_rows = [row for section in voice["domains"] for row in section["rows"]]
assert {row["capability"] for row in voice_rows} == {
"voice-gen.speech-2-8-hd.generate",
"voice-gen.speech-2-8-turbo.generate",
}
voice_endpoints = [endpoint for row in voice_rows for endpoint in row["endpoints"]]
assert {endpoint["id"] for endpoint in voice_endpoints} == {
"minimax.voice-gen.speech-2-8-hd",
"minimax.voice-gen.speech-2-8-turbo",
}
assert all(endpoint["provider"] == "minimax" for endpoint in voice_endpoints)
catalog = cs.load()
assert all(catalog.by_id[endpoint["id"]]["cache"] == "forbidden"
for endpoint in voice_endpoints)
voice_full = (await clients.get(
"/catalog/platforms/voice-gen?include_hidden=1")).json()
assert voice_full["hidden_count"] == 1
action_endpoints = {
endpoint["id"]: endpoint
for section in voice_full["domains"]
for row in section["rows"]
for endpoint in row["endpoints"]
if endpoint["kind"] == "utility"
}
assert set(action_endpoints) == {"minimax.voice-gen.voices.list"}
assert catalog.by_id["minimax.voice-gen.voices.list"]["platform_request"] == {
"body.voice_type": "system"
}
def test_a_missing_catalog_directory_is_an_empty_catalog_not_a_crash(tmp_path):
cat = cs.load(directory=tmp_path / "nope")
+6 -1
View File
@@ -190,6 +190,10 @@ def test_the_category_list_comes_from_the_data_not_the_order_list():
assert gone not in body, f"{gone} is no longer a catalog category — remove the reference"
def test_ai_generation_shelf_names_all_three_media_modalities():
assert "'AI generation':'video, image and voice models" in INDEX
def test_tiles_are_grouped_by_category_on_every_capability_tab():
""""All" is not a flat wall of tiles: it keeps the category headings, and a category tab is the
same grouping filtered to one — so the page never loses its place."""
@@ -309,7 +313,8 @@ def test_the_card_price_is_the_servers_computed_usd():
would drift from the CLI the moment the rate table changed."""
fn = INDEX[INDEX.index("platPrice(pl){") :][:1400]
assert "typeof pf.usd==='number'" in fn
assert "'$'+this.usdNum(pf.usd)+' / '+this.priceUnit(pf.type)" in fn
assert "typeof pf.display_usd==='number' ? pf.display_usd : pf.usd" in fn
assert "pf.display_unit || this.priceUnit(pf.type)" in fn
assert "return null; }" in fn # priced but unpublished → say nothing, not "from —"
# "from" is a floor: an OAuth provider among the platform's providers makes the floor $0, even
# when metered providers publish a rate — that rate demotes to the tooltip.
+1
View File
@@ -53,6 +53,7 @@ def test_key_providers_appear_in_the_marketplace_listing():
assert listing["zerobounce"]["category"] == "Enrichment"
assert listing["zerobounce"]["auth_kind"] == "key"
assert listing["minimax"]["category"] == "AI generation"
assert listing["minimax"]["summary"] == "Generate voice, images, and videos from text or source images."
assert listing["openrouter"]["auth_kind"] == "token"
assert listing["replicate"]["base_url"] == "https://api.replicate.com/v1"
assert "Enrichment" in P.CATEGORY_ORDER
+11
View File
@@ -1377,6 +1377,17 @@ def test_platform_estimate_normalizes_per_result_pricing():
assert call_resolution._platform_estimate_micro({"type": "per_call", "usd": 0.0000005}, {}) == 1
def test_platform_estimate_prices_text_to_speech_by_input_characters():
"""MiniMax publishes TTS per character, so the request's text length—not a result-page
default or a flat call price—sets the reserve. Unicode code points count as characters."""
hd = {"type": "per_success", "unit": "character", "usd": 0.0001}
turbo = {"type": "per_success", "unit": "character", "usd": 0.00006}
assert call_resolution._platform_estimate_micro(hd, {}, b'{"text":"Hello."}') == 600
assert call_resolution._platform_estimate_micro(turbo, {}, ' {"text":"Hi 👋"}'.encode()) == 240
assert call_resolution._platform_estimate_micro(hd, {}, b'{"text":""}') == 100
assert call_resolution._platform_estimate_micro(hd, {}, b'not-json') == 100
def test_platform_estimate_counts_input_entities_not_a_page():
"""A price per TARGET / DOMAIN / KEYWORD is per thing asked about, never per returned row: with
no limit param the 20-row page default billed a one-target SE Ranking summary 20x ($0.358 for a