We type a prompt and the MiniMax H3 Max API returns a finished clip with sound. The fifteen-second option tells a whole story in one render.
The MiniMax H3 Max API generates long-form video from a text prompt alone, no image needed -- 5 to 15 seconds with synchronized audio, 480P or 768P, ratio 16:9/4:3/1:1/3:4/9:16/21:9, prompts up to 7000 characters -- via one OpenAI-compatible REST endpoint, metered per clip at 15% below official.
SampleRun the model to replace this previewMiniMax H3 Max API: Text-to-Video
Use the MiniMax H3 Max API to generate long-form video from text alone: 5 to 15 seconds, synchronized audio, 480P or 768P, six aspect ratios.
The MiniMax H3 Max API (minimax/minimax-h3-max-t2v) is the long-form tier of MiniMax's H3 family: you supply only a text prompt and the MiniMax H3 Max API generates the whole clip from scratch, audio included, with no image or reference input. Duration is any whole number from 5 to 15 seconds (default 6), resolution is a choice of 480P for cheap drafts or 768P for delivery (default 768P), and ratio covers six options -- 16:9 (default), 4:3, 1:1, 3:4, 9:16 or 21:9 (text-to-video does not accept adaptive). Prompts run to 7000 characters, so the MiniMax H3 Max API takes detailed scene direction in a single request.
Pricing on the MiniMax H3 Max API is metered per clip by duration and resolution at 15% below the official list -- from $0.198 for a 5-second 480P clip, $0.360 for the default 6-second 768P clip, up to $0.900 at 15 seconds 768P -- with no per-request fee. The MiniMax H3 Max API is async and OpenAI-compatible on one RouterBase key across 200+ models: POST /v1/videos/generations with model=minimax/minimax-h3-max-t2v and a prompt, poll GET /v1/videos/generations/{id} until status=success, and pull the finished clip from a signed CDN URL.
Six reasons teams ship on the MiniMax H3 Max API
From a bare prompt to a 15-second clip with audio -- what the MiniMax H3 Max API gets right.
Generate long-form video from text
Describe the scene and the MiniMax H3 Max API generates the clip from scratch -- no image, no reference, no footage, just a prompt up to 7000 characters.
Any length from 5 to 15 seconds
Set duration to any whole second from 5 to 15 (default 6). The MiniMax H3 Max API lets one render fit a tight teaser or a full fifteen-second story without stitching.
Synchronized audio built in
Every clip from the MiniMax H3 Max API renders with synchronized audio, so a text prompt comes back as a finished, sound-on clip with no separate audio pass.
Two resolutions, one model
Pick 480P to iterate cheaply or 768P for delivery (default 768P). The MiniMax H3 Max API lets you draft and finalize on the same endpoint by flipping one parameter.
Six aspect ratios
Pick 16:9, 4:3, 1:1, 3:4, 9:16 or 21:9 per call (default 16:9). The MiniMax H3 Max API turns one prompt into a landscape, portrait, square or ultrawide cut -- text-to-video does not accept adaptive.
Metered pricing, 15% below official
From $0.198 at 5s/480P to $0.900 at 15s/768P, metered per clip with no per-request fee -- 15% off the official rate. The MiniMax H3 Max API uses the same JSON shape as every RouterBase model; switching is only a model ID change.

Get started with the MiniMax H3 Max API in 3 steps
From sign-up to your first long-form text-to-video call in under 5 minutes.
Write a descriptive prompt
Describe subject, motion, mood, sound and camera in the prompt field (up to 7000 characters). The MiniMax H3 Max API generates from text alone, so a concrete, specific prompt yields the cleanest result. Choose a duration, resolution and ratio for your placement.
POST to the MiniMax H3 Max API endpoint
Send prompt plus optional duration, resolution and ratio to POST /v1/videos/generations. The MiniMax H3 Max API returns an async job id; poll GET /v1/videos/generations/{id} until status=success.
Poll and download the clip
The success response from the MiniMax H3 Max API contains a signed CDN URL in the results array. Drop it into a <video> tag or copy it to your bucket -- audio included, no post-processing.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
| Resolution | RouterBase | Official | Save |
|---|---|---|---|
| 768P 6s | $0.360 | −15% | |
| 480P 6s | $0.238 | −15% |
About $0.060 per second.One run at the default settings (768P) costs $0.360.With $1 you can run this model approximately 2 times.
What teams build with the MiniMax H3 Max API
Real production loads running on the MiniMax H3 Max API and the RouterBase catalog.
One line of copy, a 768P clip with audio -- the MiniMax H3 Max API gives us landscape, portrait and ultrawide cuts from the same prompt.
Any length from five to fifteen seconds, no stitching. The MiniMax H3 Max API gives us exactly the duration we ask for.
We draft at 480P overnight on the MiniMax H3 Max API and re-render the winners at 768P. Same endpoint, one parameter.
Text in, sound-on video out. I sent one prompt to the MiniMax H3 Max API and had a finished clip in a single call.
RouterBase puts 200+ models behind one key. We made the MiniMax H3 Max API our default for long-form text-to-video.
At 15% below list the MiniMax H3 Max API fits our ad budget, and the six aspect ratios cover every feed.
Same JSON shape, different model ID. Adding the MiniMax H3 Max API was the easiest change in our rollout.
Metered per clip with audio included keeps costs predictable when the MiniMax H3 Max API generates thousands of clips a week.
Frequently Asked Questions
Common questions about the MiniMax H3 Max API text-to-video task.
The MiniMax H3 Max API text-to-video task generates a long-form video with synchronized audio from a text prompt alone on MiniMax's H3 Max model. No image or reference input is needed -- you describe the scene and get a finished, sound-on clip.