The DeepSeek V4 Flash API gave our assistant new-generation quality at a speed our users actually feel.
DeepSeek V4 Flash is the fast, low-cost tier of the newest DeepSeek generation — tuned for low latency and high volume, with strong coding and reliable tool calls. Served on RouterBase over an OpenAI-compatible endpoint at $0.14 / 1M input and $0.28 / 1M output, with cached input at just $0.0028; step up to DeepSeek V4 Pro when a job needs the top of the line.
DeepSeek V4 Flash API — The Fast, Low-Cost Tier of DeepSeek V4
The DeepSeek V4 Flash API serves DeepSeek V4 Flash, the speed-and-cost tier of the newest DeepSeek generation — built for latency-sensitive, high-volume work.
DeepSeek V4 Flash is the Flash tier of the newest DeepSeek generation: tuned for low latency and low cost so a capable model can sit in real-time paths and high-volume batches. The DeepSeek V4 Flash API keeps replies fast and the bill small — $0.14 per 1M input tokens and $0.28 per 1M output, with cached input at just $0.0028 — while DeepSeek V4 Pro stays a call away when a job needs the top of the line. Strong coding and reliable tool calls make it a practical default for assistants and pipelines.
You reach the DeepSeek V4 Flash API through RouterBase on the OpenAI chat-completions protocol; set the model to deepseek-v4-flash and your client is unchanged. The DeepSeek V4 Flash API bills $0.14 per 1M input, $0.28 per 1M output, and $0.0028 for cached input — 5% under list, on one key that reaches the whole catalog.
What the DeepSeek V4 Flash API gives you
A capable model that keeps up with traffic, on a budget.
New-generation Flash tier
The DeepSeek V4 Flash API is the fast, low-cost tier of the newest DeepSeek generation.
Built for speed
Low latency lets the DeepSeek V4 Flash API sit inside live requests instead of background jobs.
Small bill
At $0.14 in and $0.28 out, the DeepSeek V4 Flash API stays cheap enough for high volume — 5% below list.
Near-free cached reads
Cached input is just $0.0028 per 1M, so repeated context on the DeepSeek V4 Flash API costs almost nothing.
Coding and tools
Strong code plus reliable function calling make the DeepSeek V4 Flash API a practical backbone for assistants and agents.
Pair with V4 Pro
Route everyday work to the DeepSeek V4 Flash API and escalate to DeepSeek V4 Pro when a job needs the top of the line.

First call to the DeepSeek V4 Flash API in three steps
Quick to wire, fast to answer.
Create a key
One RouterBase key reaches the DeepSeek V4 Flash API and every other model in the catalog.
Point and name
Send any OpenAI client to routerbase.com/v1 with model=deepseek/deepseek-v4-flash. Streaming and function calling work out of the box.
Watch latency and cost
The DeepSeek V4 Flash API returns usage per response, so throughput and spend stay visible on every call.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
Teams building on the DeepSeek V4 Flash API
New-generation quality that keeps pace with the traffic.
At $0.14 in, the DeepSeek V4 Flash API let us put a capable model on every request without watching the bill.
Cached reads at $0.0028 mean our repeated context on the DeepSeek V4 Flash API is basically free.
We run everyday traffic on the DeepSeek V4 Flash API and escalate to V4 Pro only when we must.
Low latency let us drop the loading state on the DeepSeek V4 Flash API.
Function calling is reliable, so our whole tool layer runs on the DeepSeek V4 Flash API.
One string changed from our old model; the DeepSeek V4 Flash API was faster and cheaper the same day.
The DeepSeek V4 Flash API keeps our p95 low even at peak, which is exactly why we picked it.
We moved our default chat to the DeepSeek V4 Flash API and users just noticed it got faster.
DeepSeek V4 Flash API — common questions
What to know before you route traffic to it.
RouterBase's hosted access to DeepSeek V4 Flash, the fast, low-cost tier of the newest DeepSeek generation, over an OpenAI-compatible endpoint.