video_generate.vidu_q1_r2v_novita
vidu/vidu-q1/reference-to-video

The Vidu Q1 API generates a 5-second 1080p video from reference images and an optional natural-language prompt (Vidu Q1) — preserving identity and character consistency across frames, adjustable motion amplitude (auto / small / medium / large), optional background music, reproducible seed, and a single OpenAI-compatible REST endpoint, 15% below the official rate.

Input
The text prompt sent to the model
Output
Sample edited image — run the model to replace this previewSampleRun the model to replace this preview

Vidu Q1 API: Reference-to-Video

Use the Vidu Q1 API to generate consistent 5-second 1080p clips from reference images of your subject.

The Vidu Q1 API (video_generate.vidu_q1_r2v_novita) generates 5-second 1080p video clips from reference images and an optional natural-language prompt. Each Vidu Q1 API call locks the identity and appearance of your subject using the supplied reference images, producing a temporally consistent output with adjustable motion amplitude — auto, small, medium, or large — and optional background music, all through a single OpenAI-compatible REST call on the Vidu Q1 API. The Vidu Q1 API is built for character-driven work — an avatar, a mascot, or a game hero — where the same face and outfit must survive every frame, and the Vidu Q1 API reads the reference set to lock identity before rendering.

Pricing on the Vidu Q1 API is $0.34 per generation — 15% below the published rate. The Vidu Q1 API uses one RouterBase key; no Vidu-specific SDK or separate account is required. Because the Vidu Q1 API is async and OpenAI-compatible, you can queue reference-to-video jobs beside your other generations, batch a whole cast overnight, and pull finished clips from one CDN URL.

Why this model

Six reasons teams ship on the Vidu Q1 API

From reference-image consistency to per-call motion control — what the Vidu Q1 API gets right.

Identity consistency

Supply reference images and the Vidu Q1 API preserves the face, outfit, and proportions across every frame of the 5-second 1080p clip, so a character stays on-model from shot to shot.

5-second 1080p clips

Every Vidu Q1 API call produces a 5-second 1080p video — the fixed duration removes encoding guesswork and keeps cost and storage predictable across a batch.

Motion amplitude control

Pass auto, small, medium, or large per call. Small keeps a portrait steady, large drives action, and auto lets the Vidu Q1 API pick the movement from the scene.

Optional background music

Set bgm=true and the Vidu Q1 API scores the clip with a fitting track; omit the flag for a silent output you can pair with your own audio.

$0.34 / generation

15% below the published rate on the Vidu Q1 API — flat per generation regardless of motion amplitude, reference count, or BGM, so spend stays predictable.

OpenAI-compatible REST

One endpoint, one key — the Vidu Q1 API uses the same JSON shape and Bearer header as every RouterBase model, so switching to the Vidu Q1 API is only a model ID change.

RouterBase dashboard preview
Quickstart

Get started with the Vidu Q1 API in 3 steps

From sign-up to your first reference-to-video call in under 5 minutes.

  1. Upload your reference image

    Host your subject image on a public URL — the Vidu Q1 API accepts one or more reference image URLs, and extra angles sharpen identity consistency.

  2. POST to the Vidu Q1 endpoint

    Send reference_images, an optional prompt, motion_amplitude, and bgm in JSON to POST /v1/images/generations. The Vidu Q1 API returns an async job id immediately; poll GET /v1/images/generations/{id} until status=success.

  3. Download the clip

    The success response from the Vidu Q1 API contains a signed CDN URL under the results array. Drop it into a <video> tag or copy the file to your own bucket — the 5-second 1080p MP4 is ready with no post-processing.

Pricing

Pay only for what you use

RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.

ResolutionRouterBaseOfficialSave
generation $0.340$0.400−15%

One run at the default settings (generation) costs $0.340.With $1 you can run this model approximately 2 times.

Customer stories

What teams build with the Vidu Q1 API

Real production loads running on the Vidu Q1 API and the RouterBase catalog.

Marcus Reyes
Marcus ReyesCTO, Paradigm AI

We plugged reference images into the Vidu Q1 API and our avatar kept its face across every clip. Consistency we could never get before.

Priya Lakshmi
Priya LakshmiFounder, Quillo

One call with a reference photo — same character, 5-second clip, live in 90 seconds. Our influencer team loves it.

Thomas Beck
Thomas BeckStaff Engineer, Northbeam

Reference-to-video means we stopped re-uploading frames for continuity. One photo, done.

Aoi Tanaka
Aoi TanakaML Lead, Daybreak Robotics

We pipe reference images into the Vidu Q1 API nightly. BGM on, motion auto — character intact every morning.

Jonas Keller
Jonas KellerIndie Developer

Swapped the base URL to RouterBase and passed reference_images. The Vidu Q1 API kept my game character consistent across ten clips.

Ethan Nguyen
Ethan NguyenHead of Engineering, Compound Studio

RouterBase puts 200+ models behind one key. The Vidu Q1 API is the one our brand team books every sprint.

Sophia Martín
Sophia MartínCTO, Relay

At $0.34 per clip it fits inside our ad budget — and our spokesperson looks the same in every video.

Lucas Fernandes
Lucas FernandesEngineering Manager, Light

Same JSON shape, different model ID. Reference-to-video was the easiest migration in our 42-service rollout.

David Okonkwo
David OkonkwoCo-founder, Figment

15% off the rate adds up fast when you generate thousands of brand-consistent clips per week.

Frequently Asked Questions

Common questions about the Vidu Q1 API reference-to-video task.

The Vidu Q1 API reference-to-video task uses reference images of a subject to maintain identity and character consistency across the 5-second 1080p clip — unlike image-to-video, which animates a single input image. It is the right choice when a character must stay recognizable across many clips.