We route high-volume chat through the Gemini 3.5 Flash API — fast multimodal responses at scale, and the 5% RouterBase discount is pure margin.
The Gemini 3.5 Flash API is Google's fast, efficient multimodal model — a 1M-token context window, native input for text, images, audio, video and PDF, function calling, and structured JSON outputs, tuned for low latency and high throughput. The Gemini 3.5 Flash API is OpenAI-compatible: point any existing SDK at RouterBase and set the model to gemini-3-5-flash. Input at $1.05 / 1M tokens, output at $6.30 / 1M tokens, cache reads at $0.105 / 1M tokens — 5% below the standard rate, all through one REST endpoint.
Gemini 3.5 Flash API: Chat
Use the Gemini 3.5 Flash API to run Google's fast, efficient multimodal model with a 1M-token context window.
The Gemini 3.5 Flash API is Google's speed-optimized multimodal model — a 1M-token context window, native input for text, images, audio, video, and PDF, function calling, and structured JSON outputs, tuned for low latency and high throughput. Routed through RouterBase, it is fully OpenAI-compatible: point any existing SDK at RouterBase's base URL, set the model to gemini-3-5-flash, and your first call goes through immediately.
Pricing is $1.05 / 1M input tokens, $6.30 / 1M output tokens, and $0.105 / 1M cached-input reads — 5% below the standard rate. One RouterBase key — no Google credentials required.
Six reasons teams ship on the Gemini 3.5 Flash API
From fast multimodal inference to 5%-off pricing — what makes Gemini 3.5 Flash stand out.
1M-token context
Accepts up to 1 million tokens per request — long documents, multi-file code, or video transcripts fit in a single Gemini 3.5 Flash API call.
Native multimodal input
The Gemini 3.5 Flash API understands text, images, audio, video, and PDFs. Pass mixed media in one request to analyze, transcribe, or reason across modalities.
Fast & high-throughput
Tuned for low latency and high volume: the Gemini 3.5 Flash API delivers Flash-tier speed for chat, classification, extraction, and routing at scale.
Function calling & JSON
Define tools and request a JSON schema; the Gemini 3.5 Flash API returns structured tool calls and strictly valid JSON for reliable agent and extraction pipelines.
OpenAI-compatible endpoint
The Gemini 3.5 Flash API speaks the OpenAI chat-completions wire format. Point any OpenAI SDK at RouterBase — no Google SDK or separate credentials needed.
One key for 200+ models
The same RouterBase key that calls the Gemini 3.5 Flash API also routes to GPT-5, Claude Opus 4.8, GPT-4o mini, and 200+ other models — no per-provider credential management.

Get started with the Gemini 3.5 Flash API in 3 steps
From sign-up to your first response in under 5 minutes.
Create a RouterBase API key
Sign up and generate an API key — one key reaches the Gemini 3.5 Flash API and every other model in the catalog.
Send your first message
Point any OpenAI-compatible SDK at routerbase.com/v1 and set the model to gemini-3-5-flash. The response follows the standard chat-completions schema — streaming and multimodal input supported.
Inspect usage
Every response includes a detailed token breakdown — input, output, and cache-read tokens — so cost and cache-hit rate are visible on every call.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
What teams build with the Gemini 3.5 Flash API
Real production loads running on the Gemini 3.5 Flash API and the RouterBase model catalog.
Native vision on the Gemini 3.5 Flash API replaced an OCR pipeline — one fast call captions and extracts, text and image together.
Structured JSON outputs from the Gemini 3.5 Flash API killed our flaky parsing — strictly valid JSON every time, even on multimodal inputs.
Flash-tier latency on the Gemini 3.5 Flash API let us move classification online. RouterBase makes it 5% cheaper still.
The 1M context on the Gemini 3.5 Flash API means I can throw whole docs at it and still get fast answers. No chunking glue.
RouterBase puts the Gemini 3.5 Flash API and 200+ other models behind one key. Our team stopped filing requests for new provider accounts.
We A/B the Gemini 3.5 Flash API against Pro per request — same SDK, one model field. Most traffic stays on Flash and latency dropped.
We migrated 42 services to the Gemini 3.5 Flash API over a weekend. The only PR comment was 'wait, that's all?'.
Cache reads at $0.105 / 1M on the Gemini 3.5 Flash API made our long-context features affordable. The ROI math was instant.
Frequently Asked Questions
Common questions about the Gemini 3.5 Flash API.
It is RouterBase's pass-through to Google's Gemini 3.5 Flash — a fast, efficient multimodal chat endpoint with a 1M-token context, native vision/audio/video input, function calling, and structured outputs, served via an OpenAI-compatible REST interface.