We piloted the DeepSeek V3.2 Exp API to see sparse attention on our own data before committing.
Loading model information...
DeepSeek V3.2 Exp API — The Experimental Sparse-Attention Preview
The DeepSeek V3.2 Exp API serves DeepSeek-V3.2-Exp, the experimental release where DeepSeek Sparse Attention debuted — a preview of the next-gen architecture at the same low price.
DeepSeek-V3.2-Exp is the experimental model that introduced DeepSeek Sparse Attention, built on the V3.1-Terminus foundation as an intermediate step toward the next-generation architecture. The DeepSeek V3.2 Exp API lets you test that new attention on real workloads early: near-V3.1-Terminus quality with long-context cost cut sharply, while keeping the hybrid design that lets you think step by step or answer directly per request. It is the place to validate sparse attention before it becomes your default.
You call the DeepSeek V3.2 Exp API through RouterBase on the OpenAI chat-completions protocol; set the model to deepseek-v3-2-exp and your client is unchanged. The DeepSeek V3.2 Exp API bills $0.028 per 1M input, $0.42 per 1M output, and $0.028 for cached input — 5% under list, on one key that reaches the whole catalog.
What the DeepSeek V3.2 Exp API gives you
Where sparse attention first shipped.
Experimental preview
The DeepSeek V3.2 Exp API is the experimental drop where sparse attention debuted — a preview of the next-gen architecture.
Sparse attention, first
DeepSeek Sparse Attention landed here first, so the DeepSeek V3.2 Exp API cuts long-context cost sharply.
Built on V3.1-Terminus
Starting from the V3.1-Terminus foundation, the DeepSeek V3.2 Exp API holds near-parity quality while it tries the new attention.
Hybrid thinking kept
Think step by step or answer directly per request; the DeepSeek V3.2 Exp API carries both modes in one model.
Same low price
At $0.028 in and $0.42 out, the DeepSeek V3.2 Exp API costs the same as the stable line — 5% below list.
OpenAI-shaped
One chat-completions endpoint, one RouterBase key — point at routerbase.com/v1, name the model, skip the DeepSeek SDK.

First call to the DeepSeek V3.2 Exp API in three steps
Wire it once, then compare it to your current model.
Create a key
One RouterBase key reaches the DeepSeek V3.2 Exp API and every other model in the catalog.
Point and pick a mode
Send any OpenAI client to routerbase.com/v1 with model=deepseek/deepseek-v3.2-exp, and toggle thinking on or off per request.
Test and compare
The DeepSeek V3.2 Exp API returns token usage per response, so you can weigh quality and cost against your current model.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
Teams testing the DeepSeek V3.2 Exp API
Sparse attention, tried on real workloads.
Near-Terminus quality at a fraction of the long-context cost — the DeepSeek V3.2 Exp API was worth the experiment.
Running the DeepSeek V3.2 Exp API next to our stable model made the comparison easy.
At $0.028 in, testing the DeepSeek V3.2 Exp API cost us almost nothing.
We flipped one string and had the DeepSeek V3.2 Exp API answering our long-context prompts.
The hybrid modes carried over, so our eval harness ran on the DeepSeek V3.2 Exp API unchanged.
Sparse attention held up on our longest documents when we tried the DeepSeek V3.2 Exp API.
The DeepSeek V3.2 Exp API let us preview the next-gen architecture without a migration.
We kept our stable model in production and used the DeepSeek V3.2 Exp API to plan the switch.
DeepSeek V3.2 Exp API — common questions
What to know before you pilot it.
RouterBase's hosted access to DeepSeek-V3.2-Exp, the experimental release that introduced DeepSeek Sparse Attention, over an OpenAI-compatible endpoint.