Our writers type a prompt and the Vidu Q2 API returns a clip. No image pipeline, no upload step -- text goes in, video comes out.
The Vidu Q2 API generates video from a text prompt alone, no image needed — variable 1-10s duration, up to 1080p, three aspect ratios (16:9/9:16/1:1), optional style hint, BGM and audio, reproducible seed — via one OpenAI-compatible REST endpoint 15% below official.
SampleRun the model to replace this previewVidu Q2 API: Text-to-Video
Turn a plain text prompt into video with the Vidu Q2 API — no image, up to 1080p, 1-10 seconds.
The Vidu Q2 API (video_generate.vidu_q2_t2v_novita) is a text-to-video model: the Vidu Q2 API generates video from a natural-language prompt alone, with no image and no reference input. You describe the scene in words and the Vidu Q2 API renders a clip up to 1080p at a variable 1 to 10 second duration. Aspect ratio (16:9, 9:16, 1:1), motion amplitude, an optional style hint, optional audio, and optional background music are all controllable through one OpenAI-compatible REST call on the Vidu Q2 API.
Pricing on the Vidu Q2 API is $0.149 per generation — 15% below the official rate of $0.175. That economical Q2 tier makes the Vidu Q2 API practical for cheap iteration: draft a dozen variations, compare aspect ratios, and lock a seed before a final render. It uses one RouterBase key, no separate SDK or account, and the same JSON shape covers 200+ other models. Because the Vidu Q2 API is async and OpenAI-compatible, you queue text-to-video jobs, batch a campaign overnight, and pull finished clips from a signed CDN URL, with any seed making the Vidu Q2 API reproducible.
Six reasons teams ship on the Vidu Q2 API
From text-only input to 1080p output — what the Vidu Q2 API gets right.
Text-only input
Send a prompt and nothing else — the Vidu Q2 API needs no image, no reference, and no upload, so an idea becomes a clip the moment you finish typing the sentence.
Up to 1080p resolution
Render at 540p, 720p, or up to 1080p — the Vidu Q2 API lets you pick resolution per call, so you draft cheaply at 540p and ship the final at 1080p.
1-10 second variable duration
Set any duration from 1 to 10 seconds. The Vidu Q2 API is not locked to a fixed clip length, so one render fits a short loop or a longer social cut without padding.
Three aspect ratios & motion
Choose 16:9, 9:16, or 1:1 and dial motion amplitude (auto, small, medium, high) in a single Vidu Q2 API request — one call per platform, no re-editing.
Style hint, audio & BGM
Steer the look with an optional style hint, then toggle audio and background music on the Vidu Q2 API to ship a clip that sounds finished, not silent.
OpenAI-compatible & 15% off
One endpoint, one key at $0.149/generation — 15% below the official rate, and the Vidu Q2 API uses the same JSON shape as every RouterBase model, so adopting it is only a model ID change.

Get started with the Vidu Q2 API in 3 steps
From sign-up to your first text-to-video call in under 5 minutes.
Write your prompt
Describe the scene, subject, and mood in plain language, and add an optional style hint — the Vidu Q2 API is text-only, so a vivid, specific sentence is the whole input.
POST to the Vidu Q2 endpoint
Send prompt and options like duration, resolution, aspect_ratio, and seed in JSON to POST /v1/images/generations. The Vidu Q2 API returns an async job id immediately; poll GET /v1/images/generations/{id} until status=success.
Poll and download the clip
The success response from the Vidu Q2 API contains a signed CDN URL under the results array. Drop it into a <video> tag or copy the file to your own bucket — it is ready with no post-processing.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
| Resolution | RouterBase | Official | Save |
|---|---|---|---|
| 540p 5s | $0.076 | −15% | |
| 720p 5s | $0.149 | −15% | |
| 1080p 5s | $0.255 | −15% |
About $0.030 per second.One run at the default settings (720p) costs $0.149.With $1 you can run this model approximately 6 times.
What teams build with the Vidu Q2 API
Real production loads running on the Vidu Q2 API and the RouterBase catalog.
One sentence, a 9:16 clip, done. Text-to-video is how our social team ships shorts in minutes now.
Variable 1-10s duration means we do not pad or trim anymore. The Vidu Q2 API gives us exactly the length we ask for.
We render nightly at 1080p with BGM on, straight from prompts. Every clip is ready by morning.
I passed a style hint and a prompt to the Vidu Q2 API and my trailer shots matched the game art immediately.
RouterBase puts 200+ models behind one key. We made text-to-video our marketing team default.
At $0.149 per clip we iterate a dozen variations before we commit. Cheap drafts changed how we work.
Same JSON shape, different model ID. This was the easiest addition in our 42-service rollout.
15% off the rate adds up fast when a prompt-only workflow lets us generate thousands of clips a week.
Frequently Asked Questions
Common questions about the Vidu Q2 API text-to-video task.
The Vidu Q2 API text-to-video task generates a clip from a text prompt alone — no image and no reference input. You describe the scene in words and the model renders the video, which makes it the right choice for idea-to-video work and high-volume social clips.