chat.claude_haiku_4_5
anthropic/claude-haiku-4-5

The Claude Haiku 4.5 API is Anthropic's fastest, most cost-effective model — near-frontier intelligence at small-model speed and price, with a 200K-token context window, extended thinking, vision, tool use, and prompt caching. The Claude Haiku 4.5 API is OpenAI-compatible: point any existing SDK at RouterBase and swap the model name to claude-haiku-4-5. Input at $0.95 / 1M tokens, output at $4.75 / 1M tokens, cache reads at $0.095 / 1M tokens — 5% below the standard rate, all through one REST endpoint.

Input
Claude Haiku 4.5
max_tokens4096
temperature1
top_p1
presence_penalty0
frequency_penalty0
Claude Haiku 4.5 Online
Hi! I'm a helpful AI assistant. What can I do for you?

Claude Haiku 4.5 API: Chat

Use the Claude Haiku 4.5 API to run Anthropic's fastest, most cost-effective model — a 200K-token context, extended thinking, and prompt caching.

The Claude Haiku 4.5 API is Anthropic's fastest, most cost-effective model — near-frontier intelligence at small-model speed and price, with a 200K-token context window, extended thinking, vision, tool use, and prompt caching. Routed through RouterBase, the Claude Haiku 4.5 API is fully OpenAI-compatible: point any existing SDK at RouterBase's base URL, swap the model name to claude-haiku-4-5, and your first call goes through immediately.

Pricing is $0.95 / 1M input tokens, $4.75 / 1M output tokens, and $0.095 / 1M cached-input reads — 5% below the standard rate. One RouterBase key — no Anthropic account required.

Why this model

Six reasons teams ship on the Claude Haiku 4.5 API

From near-frontier speed to 5%-off pricing — what makes the Claude Haiku 4.5 API stand out.

Fastest, cheapest Claude

The Claude Haiku 4.5 API delivers near-frontier quality at the lowest Claude price and the highest speed — ideal for high-volume, latency-sensitive workloads.

Extended thinking

Enable extended thinking to let the Claude Haiku 4.5 API reason step-by-step when a task needs it, or keep it off for the fastest, cheapest responses — you control the depth per request.

200K-token context

Accepts up to 200,000 tokens per request — long documents, multi-file codebases, or full conversation history fit in a single Claude Haiku 4.5 API call without chunking.

Vision & tool use

The Claude Haiku 4.5 API accepts images and returns structured tool calls in the standard OpenAI schema — well-suited to extraction, classification, and fast agent steps.

Prompt caching

Cache long system prompts and shared context; the Claude Haiku 4.5 API bills cached reads at $0.095 / 1M, cutting cost up to 90% on repeated prefixes.

OpenAI-compatible endpoint

The Claude Haiku 4.5 API speaks the OpenAI chat-completions wire format. Point any OpenAI SDK at RouterBase — no Anthropic SDK or separate credentials needed.

RouterBase dashboard preview
Quickstart

Get started with the Claude Haiku 4.5 API in 3 steps

From sign-up to your first response in under 5 minutes.

  1. Create a RouterBase API key

    Sign up and generate an API key — one key reaches the Claude Haiku 4.5 API and every other model in the catalog.

  2. Send your first message

    Point any OpenAI-compatible SDK at routerbase.com/v1 and set the model to claude-haiku-4-5. The response follows the standard chat-completions schema — streaming and tool use supported.

  3. Inspect usage

    Every response includes a detailed token breakdown — input, output, and cache-read tokens — so cost and cache-hit rate are visible on every call.

Pricing

Pay only for what you use

RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.

Customer stories

What teams build with the Claude Haiku 4.5 API

Real production loads running on the Claude Haiku 4.5 API and the RouterBase model catalog.

Marcus Reyes
Marcus ReyesCTO, Paradigm AI

The Claude Haiku 4.5 API runs our highest-volume traffic — fast, cheap, and good enough that most calls never need a bigger model. The 5% RouterBase discount is pure margin.

Priya Lakshmi
Priya LakshmiFounder, Quillo

We moved classification and routing to the Claude Haiku 4.5 API and latency dropped while quality held. It handles millions of calls a day without blinking.

Thomas Beck
Thomas BeckStaff Engineer, Northbeam

Prompt caching on the Claude Haiku 4.5 API dropped our extraction bill by more than half — the long system prompt is basically free now.

Aoi Tanaka
Aoi TanakaML Lead, Daybreak Robotics

Extended thinking on the Claude Haiku 4.5 API lets us turn on reasoning only when a task needs it. RouterBase makes it 5% cheaper and routes around outages automatically.

Jonas Keller
Jonas KellerIndie Developer

200K of context at Haiku prices means I can stuff whole docs in cheaply. The Claude Haiku 4.5 API answers fast and barely moves the bill.

Ethan Nguyen
Ethan NguyenHead of Engineering, Compound Studio

RouterBase puts the Claude Haiku 4.5 API and 200+ other models behind one key. Our team stopped filing requests for new provider accounts.

Sophia Martín
Sophia MartínCTO, Relay

We A/B the Claude Haiku 4.5 API against Sonnet per request — same SDK, one model field. Most traffic stays on Haiku and the bill dropped sharply.

Lucas Fernandes
Lucas FernandesEngineering Manager, Light

We routed our high-volume endpoints to the Claude Haiku 4.5 API over a weekend. The only PR comment was 'wait, that's all?'.

David Okonkwo
David OkonkwoCo-founder, Figment

At $0.95 / 1M input minus 5%, the Claude Haiku 4.5 API gave us near-frontier quality at a price that scales. The ROI math was instant.

Frequently Asked Questions

Common questions about the Claude Haiku 4.5 API.

It is RouterBase's pass-through to Anthropic's Claude Haiku 4.5 — Anthropic's fastest, most cost-effective model, with a 200K-token context, extended thinking, vision, tool use, and prompt caching, served via an OpenAI-compatible REST interface.