video_generate.vidu_q2_r2v_novita
vidu/vidu-q2/reference-to-video

The Vidu Q2 API generates video from reference subjects and a text prompt (Vidu Q2), preserving subject identity across frames — variable 1-10 second duration, up to 1080p, aspect-ratio control (16:9 / 9:16 / 1:1), adjustable motion amplitude, optional background music and audio, and a reproducible seed, via one OpenAI-compatible REST endpoint at 15% below the official rate.

Input
The text prompt sent to the model
Output
Sample edited image — run the model to replace this previewSampleRun the model to replace this preview

Vidu Q2 API: Reference-to-Video

Use the Vidu Q2 API to generate consistent video from reference subjects — up to 1080p, 1-10 seconds.

The Vidu Q2 API (video_generate.vidu_q2_r2v_novita) generates video from reference subjects — people, characters, or objects — and a natural-language prompt. Each Vidu Q2 API call preserves the identity and appearance of your subjects across every frame, with a variable 1-10 second duration and resolution up to 1080p. Aspect ratio (16:9 / 9:16 / 1:1), motion amplitude, optional background music, and audio are all controllable through a single OpenAI-compatible REST call on the Vidu Q2 API. Teams reach for the Vidu Q2 API when a recurring character, mascot, or product must stay recognizable across many clips, and the variable duration lets one render fit a six-second ad or a ten-second cut without trimming.

Pricing on the Vidu Q2 API is $0.191 per generation — 15% below the official rate of $0.225. The Vidu Q2 API uses one RouterBase key; no separate SDK or account is required, and the same JSON shape covers 200+ other models. Because the Vidu Q2 API is async and OpenAI-compatible, you can queue reference-to-video jobs beside your other generations, batch a whole cast overnight, and pull finished clips from a signed CDN URL.

Why this model

Six reasons teams ship on the Vidu Q2 API

From reference-subject consistency to 1080p output — what the Vidu Q2 API gets right.

Reference-subject consistency

Supply reference subjects and the Vidu Q2 API keeps their face, outfit, and proportions stable across every frame, so a character stays on-model from shot to shot.

Up to 1080p resolution

Render at 540p, 720p, or up to 1080p — the Vidu Q2 API lets you pick resolution per call, so you can draft cheaply at 540p and ship the final at 1080p.

1-10 second variable duration

Set any duration from 1 to 10 seconds. The Vidu Q2 API is not locked to a fixed clip length, so one render fits a short loop or a longer social cut without padding.

Aspect ratio & motion

Choose 16:9, 9:16, or 1:1 and dial motion amplitude (auto, small, medium, high) in a single Vidu Q2 API request — one call per platform, no re-editing.

Optional BGM, audio & seed

Toggle background music and audio on the Vidu Q2 API, and pass an integer seed for reproducible results across retries and approvals.

OpenAI-compatible & 15% off

One endpoint, one key at $0.191/generation — 15% below the official rate, and the Vidu Q2 API uses the same JSON shape as every RouterBase model, so switching to it is only a model ID change.

RouterBase dashboard preview
Quickstart

Get started with the Vidu Q2 API in 3 steps

From sign-up to your first reference-to-video call in under 5 minutes.

  1. Define your reference subjects

    Describe the people, characters, or objects to preserve and supply them as subjects — the Vidu Q2 API uses them to lock identity, and clean, well-lit references sharpen consistency.

  2. POST to the Vidu Q2 endpoint

    Send prompt, subjects, and options like duration, resolution, and aspect_ratio in JSON to POST /v1/images/generations. The Vidu Q2 API returns an async job id immediately; poll GET /v1/images/generations/{id} until status=success.

  3. Poll and download the clip

    The success response from the Vidu Q2 API contains a signed CDN URL under the results array. Drop it into a <video> tag or copy the file to your own bucket — it is ready with no post-processing.

Pricing

Pay only for what you use

RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.

ResolutionRouterBaseOfficialSave
540p 5s$0.149$0.175−15%
720p 5s$0.191$0.225−15%
1080p 5s$0.489$0.575−15%
audio=true +$0.064+$0.075−15%

About $0.038 per second.One run at the default settings (720p) costs $0.191.With $1 you can run this model approximately 5 times.

Customer stories

What teams build with the Vidu Q2 API

Real production loads running on the Vidu Q2 API and the RouterBase catalog.

Marcus Reyes
Marcus ReyesCTO, Paradigm AI

We fed reference subjects into the Vidu Q2 API and our characters held their identity at 1080p. It's the consistency we couldn't get before.

Priya Lakshmi
Priya LakshmiFounder, Quillo

One call, a 9:16 clip, same character — our influencer team ships in minutes now.

Thomas Beck
Thomas BeckStaff Engineer, Northbeam

Variable 1-10s duration means we don't pad or trim anymore. The Vidu Q2 API gives us exactly the length we ask for.

Aoi Tanaka
Aoi TanakaML Lead, Daybreak Robotics

We render nightly at 1080p with BGM on. Every subject stays intact by morning.

Jonas Keller
Jonas KellerIndie Developer

I passed my subjects and prompt to the Vidu Q2 API and my game character stayed consistent across ten clips.

Ethan Nguyen
Ethan NguyenHead of Engineering, Compound Studio

RouterBase puts 200+ models behind one key. We've made reference-to-video our brand team's default.

Sophia Martín
Sophia MartínCTO, Relay

At $0.191 per clip it fits our ad budget, and aspect-ratio control means one render per platform.

Lucas Fernandes
Lucas FernandesEngineering Manager, Light

Same JSON shape, different model ID. This was the easiest addition in our 42-service rollout.

David Okonkwo
David OkonkwoCo-founder, Figment

15% off the rate adds up fast when you generate thousands of brand-consistent clips per week.

Frequently Asked Questions

Common questions about the Vidu Q2 API reference-to-video task.

The Vidu Q2 API reference-to-video task uses reference subjects such as people, characters, or objects to keep identity and appearance consistent across the generated clip. It is the right choice when the same subject must stay recognizable across a series of videos.