The Claude Haiku 4.5 API runs our highest-volume traffic — fast, cheap, and good enough that most calls never need a bigger model. The 5% RouterBase discount is pure margin.
The Claude Haiku 4.5 API is Anthropic's fastest, most cost-effective model — near-frontier intelligence at small-model speed and price, with a 200K-token context window, extended thinking, vision, tool use, and prompt caching. The Claude Haiku 4.5 API is OpenAI-compatible: point any existing SDK at RouterBase and swap the model name to claude-haiku-4-5. Input at $0.95 / 1M tokens, output at $4.75 / 1M tokens, cache reads at $0.095 / 1M tokens — 5% below the standard rate, all through one REST endpoint.
Claude Haiku 4.5 API: Chat
Use the Claude Haiku 4.5 API to run Anthropic's fastest, most cost-effective model — a 200K-token context, extended thinking, and prompt caching.
The Claude Haiku 4.5 API is Anthropic's fastest, most cost-effective model — near-frontier intelligence at small-model speed and price, with a 200K-token context window, extended thinking, vision, tool use, and prompt caching. Routed through RouterBase, the Claude Haiku 4.5 API is fully OpenAI-compatible: point any existing SDK at RouterBase's base URL, swap the model name to claude-haiku-4-5, and your first call goes through immediately.
Pricing is $0.95 / 1M input tokens, $4.75 / 1M output tokens, and $0.095 / 1M cached-input reads — 5% below the standard rate. One RouterBase key — no Anthropic account required.
Six reasons teams ship on the Claude Haiku 4.5 API
From near-frontier speed to 5%-off pricing — what makes the Claude Haiku 4.5 API stand out.
Fastest, cheapest Claude
The Claude Haiku 4.5 API delivers near-frontier quality at the lowest Claude price and the highest speed — ideal for high-volume, latency-sensitive workloads.
Extended thinking
Enable extended thinking to let the Claude Haiku 4.5 API reason step-by-step when a task needs it, or keep it off for the fastest, cheapest responses — you control the depth per request.
200K-token context
Accepts up to 200,000 tokens per request — long documents, multi-file codebases, or full conversation history fit in a single Claude Haiku 4.5 API call without chunking.
Vision & tool use
The Claude Haiku 4.5 API accepts images and returns structured tool calls in the standard OpenAI schema — well-suited to extraction, classification, and fast agent steps.
Prompt caching
Cache long system prompts and shared context; the Claude Haiku 4.5 API bills cached reads at $0.095 / 1M, cutting cost up to 90% on repeated prefixes.
OpenAI-compatible endpoint
The Claude Haiku 4.5 API speaks the OpenAI chat-completions wire format. Point any OpenAI SDK at RouterBase — no Anthropic SDK or separate credentials needed.

Get started with the Claude Haiku 4.5 API in 3 steps
From sign-up to your first response in under 5 minutes.
Create a RouterBase API key
Sign up and generate an API key — one key reaches the Claude Haiku 4.5 API and every other model in the catalog.
Send your first message
Point any OpenAI-compatible SDK at routerbase.com/v1 and set the model to claude-haiku-4-5. The response follows the standard chat-completions schema — streaming and tool use supported.
Inspect usage
Every response includes a detailed token breakdown — input, output, and cache-read tokens — so cost and cache-hit rate are visible on every call.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
What teams build with the Claude Haiku 4.5 API
Real production loads running on the Claude Haiku 4.5 API and the RouterBase model catalog.
We moved classification and routing to the Claude Haiku 4.5 API and latency dropped while quality held. It handles millions of calls a day without blinking.
Prompt caching on the Claude Haiku 4.5 API dropped our extraction bill by more than half — the long system prompt is basically free now.
Extended thinking on the Claude Haiku 4.5 API lets us turn on reasoning only when a task needs it. RouterBase makes it 5% cheaper and routes around outages automatically.
200K of context at Haiku prices means I can stuff whole docs in cheaply. The Claude Haiku 4.5 API answers fast and barely moves the bill.
RouterBase puts the Claude Haiku 4.5 API and 200+ other models behind one key. Our team stopped filing requests for new provider accounts.
We A/B the Claude Haiku 4.5 API against Sonnet per request — same SDK, one model field. Most traffic stays on Haiku and the bill dropped sharply.
We routed our high-volume endpoints to the Claude Haiku 4.5 API over a weekend. The only PR comment was 'wait, that's all?'.
At $0.95 / 1M input minus 5%, the Claude Haiku 4.5 API gave us near-frontier quality at a price that scales. The ROI math was instant.
Frequently Asked Questions
Common questions about the Claude Haiku 4.5 API.
It is RouterBase's pass-through to Anthropic's Claude Haiku 4.5 — Anthropic's fastest, most cost-effective model, with a 200K-token context, extended thinking, vision, tool use, and prompt caching, served via an OpenAI-compatible REST interface.