video_generate.vidu_q2_t2v_novita
vidu/vidu-q2/text-to-video

The Vidu Q2 API generates video from a text prompt alone, no image needed — variable 1-10s duration, up to 1080p, three aspect ratios (16:9/9:16/1:1), optional style hint, BGM and audio, reproducible seed — via one OpenAI-compatible REST endpoint 15% below official.

Input
The text prompt sent to the model
Output
Sample edited image — run the model to replace this previewSampleRun the model to replace this preview

Vidu Q2 API: Text-to-Video

Turn a plain text prompt into video with the Vidu Q2 API — no image, up to 1080p, 1-10 seconds.

The Vidu Q2 API (video_generate.vidu_q2_t2v_novita) is a text-to-video model: the Vidu Q2 API generates video from a natural-language prompt alone, with no image and no reference input. You describe the scene in words and the Vidu Q2 API renders a clip up to 1080p at a variable 1 to 10 second duration. Aspect ratio (16:9, 9:16, 1:1), motion amplitude, an optional style hint, optional audio, and optional background music are all controllable through one OpenAI-compatible REST call on the Vidu Q2 API.

Pricing on the Vidu Q2 API is $0.149 per generation — 15% below the official rate of $0.175. That economical Q2 tier makes the Vidu Q2 API practical for cheap iteration: draft a dozen variations, compare aspect ratios, and lock a seed before a final render. It uses one RouterBase key, no separate SDK or account, and the same JSON shape covers 200+ other models. Because the Vidu Q2 API is async and OpenAI-compatible, you queue text-to-video jobs, batch a campaign overnight, and pull finished clips from a signed CDN URL, with any seed making the Vidu Q2 API reproducible.

Why this model

Six reasons teams ship on the Vidu Q2 API

From text-only input to 1080p output — what the Vidu Q2 API gets right.

Text-only input

Send a prompt and nothing else — the Vidu Q2 API needs no image, no reference, and no upload, so an idea becomes a clip the moment you finish typing the sentence.

Up to 1080p resolution

Render at 540p, 720p, or up to 1080p — the Vidu Q2 API lets you pick resolution per call, so you draft cheaply at 540p and ship the final at 1080p.

1-10 second variable duration

Set any duration from 1 to 10 seconds. The Vidu Q2 API is not locked to a fixed clip length, so one render fits a short loop or a longer social cut without padding.

Three aspect ratios & motion

Choose 16:9, 9:16, or 1:1 and dial motion amplitude (auto, small, medium, high) in a single Vidu Q2 API request — one call per platform, no re-editing.

Style hint, audio & BGM

Steer the look with an optional style hint, then toggle audio and background music on the Vidu Q2 API to ship a clip that sounds finished, not silent.

OpenAI-compatible & 15% off

One endpoint, one key at $0.149/generation — 15% below the official rate, and the Vidu Q2 API uses the same JSON shape as every RouterBase model, so adopting it is only a model ID change.

RouterBase dashboard preview
Quickstart

Get started with the Vidu Q2 API in 3 steps

From sign-up to your first text-to-video call in under 5 minutes.

  1. Write your prompt

    Describe the scene, subject, and mood in plain language, and add an optional style hint — the Vidu Q2 API is text-only, so a vivid, specific sentence is the whole input.

  2. POST to the Vidu Q2 endpoint

    Send prompt and options like duration, resolution, aspect_ratio, and seed in JSON to POST /v1/images/generations. The Vidu Q2 API returns an async job id immediately; poll GET /v1/images/generations/{id} until status=success.

  3. Poll and download the clip

    The success response from the Vidu Q2 API contains a signed CDN URL under the results array. Drop it into a <video> tag or copy the file to your own bucket — it is ready with no post-processing.

Pricing

Pay only for what you use

RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.

ResolutionRouterBaseOfficialSave
540p 5s$0.076$0.090−15%
720p 5s$0.149$0.175−15%
1080p 5s$0.255$0.300−15%

About $0.030 per second.One run at the default settings (720p) costs $0.149.With $1 you can run this model approximately 6 times.

Customer stories

What teams build with the Vidu Q2 API

Real production loads running on the Vidu Q2 API and the RouterBase catalog.

Marcus Reyes
Marcus ReyesCTO, Paradigm AI

Our writers type a prompt and the Vidu Q2 API returns a clip. No image pipeline, no upload step -- text goes in, video comes out.

Priya Lakshmi
Priya LakshmiFounder, Quillo

One sentence, a 9:16 clip, done. Text-to-video is how our social team ships shorts in minutes now.

Thomas Beck
Thomas BeckStaff Engineer, Northbeam

Variable 1-10s duration means we do not pad or trim anymore. The Vidu Q2 API gives us exactly the length we ask for.

Aoi Tanaka
Aoi TanakaML Lead, Daybreak Robotics

We render nightly at 1080p with BGM on, straight from prompts. Every clip is ready by morning.

Jonas Keller
Jonas KellerIndie Developer

I passed a style hint and a prompt to the Vidu Q2 API and my trailer shots matched the game art immediately.

Ethan Nguyen
Ethan NguyenHead of Engineering, Compound Studio

RouterBase puts 200+ models behind one key. We made text-to-video our marketing team default.

Sophia MartinCTO, Relay

At $0.149 per clip we iterate a dozen variations before we commit. Cheap drafts changed how we work.

Lucas Fernandes
Lucas FernandesEngineering Manager, Light

Same JSON shape, different model ID. This was the easiest addition in our 42-service rollout.

David Okonkwo
David OkonkwoCo-founder, Figment

15% off the rate adds up fast when a prompt-only workflow lets us generate thousands of clips a week.

Frequently Asked Questions

Common questions about the Vidu Q2 API text-to-video task.

The Vidu Q2 API text-to-video task generates a clip from a text prompt alone — no image and no reference input. You describe the scene in words and the model renders the video, which makes it the right choice for idea-to-video work and high-volume social clips.