We pointed our invoice pipeline at the Qwen3 VL 30B A3B Instruct API and the OCR plus JSON extraction replaced two separate services in a week.
Qwen3 VL 30B A3B Instruct is a Mixture-of-Experts vision-language model from Alibaba Qwen with 30B total and about 3B active parameters, reading images and text in one call with strong OCR and chart understanding. Served on RouterBase over an OpenAI-compatible endpoint at $0.17 / 1M input and $0.595 / 1M output, 15% below list.
Qwen3 VL 30B A3B Instruct API - Vision-Language MoE, 3B Active
The Qwen3 VL 30B A3B Instruct API runs a Mixture-of-Experts vision-language model with 30B total and 3B active parameters that reads images and text.
The Qwen3 VL 30B A3B Instruct API serves Alibaba Qwen3 VL 30B A3B Instruct, a Mixture-of-Experts vision-language model that activates only about 3B of its 30B parameters per token, giving big-network perception at small-model latency and cost. Screenshots, photos, scans, charts, and diagrams ride next to your prompt as standard image_url parts, so the Qwen3 VL 30B A3B Instruct API works with any OpenAI client unchanged, and one key covers the whole catalog.
Instruction-tuned for direct answers, the model handles OCR from photos and scans, document and chart understanding, screenshot analysis, and visual grounding. The Qwen3 VL 30B A3B Instruct API bills $0.17 per 1M input tokens and $0.595 per 1M output tokens, 15% below the list rate of $0.20 in and $0.70 out, with no per-request fee. Streaming and standard-schema tool calls let you wire the Qwen3 VL 30B A3B Instruct API into a document pipeline or a GUI agent without provider-specific glue.
What the Qwen3 VL 30B A3B Instruct API gives you
Vision-language quality at 3B-active cost.
Image plus text input
The Qwen3 VL 30B A3B Instruct API takes images and text in the same request, so screenshots, photos, scans, charts, and diagrams ride next to your prompt in one messages array.
MoE with 3B active
A Mixture-of-Experts network with 30B total parameters activates only about 3B per token, so the Qwen3 VL 30B A3B Instruct API answers with big-model perception at small-model latency and cost.
OCR and documents
Strong text recognition across photos, scans, and document pages makes the Qwen3 VL 30B A3B Instruct API a dependable base for receipt parsing, form reading, and document extraction.
Charts, UIs, grounding
The model reads charts and tables, understands app screenshots and GUIs, and localizes objects in a scene, so the Qwen3 VL 30B A3B Instruct API fits agents that look before they act.
Tools and JSON
The Qwen3 VL 30B A3B Instruct API returns function calls and structured JSON in the standard OpenAI schema, ready for multimodal agents and data pipelines.
Streaming, per token
The Qwen3 VL 30B A3B Instruct API streams tokens as they generate and bills per token at $0.17 in and $0.595 out, 15% below list, with no request fee.

First call to the Qwen3 VL 30B A3B Instruct API in three steps
Wire it once, then stream.
Create a key
One RouterBase key reaches the Qwen3 VL 30B A3B Instruct API and every other model in the catalog, so there is nothing provider-specific to set up.
Point your client
Send any OpenAI client to routerbase.com/v1 with model=qwen/qwen3-vl-30b-a3b-instruct and a messages array that mixes text and image_url parts; the Qwen3 VL 30B A3B Instruct API answers on the same chat-completions shape.
Stream and watch usage
The model streams tokens as they generate and returns token usage per response, so cost and latency stay visible on every call to the Qwen3 VL 30B A3B Instruct API.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
| Resolution | RouterBase | Official | Save |
|---|---|---|---|
| Input (per 1M) | $0.170 | −15% | |
| Output (per 1M) | $0.595 | −15% |
One run at the default settings (Input (per 1M)) costs $0.170.With $1 you can run this model approximately 5 times.
Official price = Alibaba Qwen list rate for Qwen3 VL 30B A3B Instruct ($0.20 per 1M input, $0.70 per 1M output). RouterBase serves the Qwen3 VL 30B A3B Instruct API at $0.17 / $0.595, 15% below list.
Teams building on the Qwen3 VL 30B A3B Instruct API
Screens, scans, and photos behind one call.
With 3B active parameters the Qwen3 VL 30B A3B Instruct API answers screenshot questions faster than our old dense VL model, at a fraction of the cost.
Standard image_url parts meant our multimodal stack ran on the Qwen3 VL 30B A3B Instruct API with zero custom plumbing on day one.
At $0.17 in and $0.595 out the Qwen3 VL 30B A3B Instruct API made per-page document analysis cheap enough to run on every upload.
Streaming from the Qwen3 VL 30B A3B Instruct API is smooth, so our chat UI renders answers about images token by token.
We send chart screenshots to the Qwen3 VL 30B A3B Instruct API and the numbers come back as clean structured JSON every time.
Our GUI agent reads app screens through the Qwen3 VL 30B A3B Instruct API and grounds the buttons it needs to click reliably.
Switching from another provider to the Qwen3 VL 30B A3B Instruct API was one string change - the OpenAI shape did not move at all.
Multilingual OCR is why we standardized on the Qwen3 VL 30B A3B Instruct API; Portuguese receipts parse as cleanly as English ones.
Qwen3 VL 30B A3B Instruct API - common questions
What to know before you route traffic to it.
The Qwen3 VL 30B A3B Instruct API is RouterBase hosted access to Alibaba Qwen3 VL 30B A3B Instruct, a Mixture-of-Experts vision-language model with 30B total and 3B active parameters, over an OpenAI-compatible endpoint.