The Gemini 3.1 Flash-Lite API is our high-volume default — newest-gen quality, fast, and cheap enough to skip the flagship on most calls. The 5% RouterBase discount is pure margin.
The Gemini 3.1 Flash-Lite API is Google's fastest, most cost-effective Gemini 3.1 model — a 1M-token context window, native input for text, images, audio, video and PDF, function calling, and structured JSON outputs, tuned for low latency and high throughput. The Gemini 3.1 Flash-Lite API is OpenAI-compatible: point any existing SDK at RouterBase and set the model to gemini-3-1-flash-lite. Input at $0.2375 / 1M tokens, output at $1.425 / 1M tokens, cache reads at $0.02375 / 1M tokens — 5% below the official published rate, all through one REST endpoint.
Gemini 3.1 Flash-Lite API: Chat
Use the Gemini 3.1 Flash-Lite API to run Google's fastest, cheapest Gemini 3.1 model — a 1M-token context, native multimodal input, and structured outputs.
The Gemini 3.1 Flash-Lite API is Google's fastest, most cost-effective Gemini 3.1 model — a 1M-token context window, native input for text, images, audio, video, and PDF, function calling, and structured JSON outputs, tuned for low latency and high throughput. Routed through RouterBase, the Gemini 3.1 Flash-Lite API is fully OpenAI-compatible: point any existing SDK at RouterBase's base URL, set the model to gemini-3-1-flash-lite, and your first call goes through immediately.
Pricing is $0.2375 / 1M input tokens, $1.425 / 1M output tokens, and $0.02375 / 1M cached-input reads — 5% below the official published rate. One RouterBase key — no Google credentials required.
Six reasons teams ship on the Gemini 3.1 Flash-Lite API
From newest-gen speed to 5%-off pricing — what makes the Gemini 3.1 Flash-Lite API stand out.
Fastest, cheapest Gemini 3.1
The Gemini 3.1 Flash-Lite API is Google’s newest lightweight model — built for the lowest cost and highest throughput in the Gemini 3.1 family.
1M-token context
Accepts up to 1 million tokens per request — long documents, multi-file code, or video transcripts fit in a single Gemini 3.1 Flash-Lite API call without chunking.
Native multimodal input
The Gemini 3.1 Flash-Lite API understands text, images, audio, video, and PDFs. Pass mixed media in one request to analyze, transcribe, or reason across modalities.
Function calling & JSON
Define tools and request a JSON schema; the Gemini 3.1 Flash-Lite API returns structured tool calls and strictly valid JSON for reliable agent and extraction pipelines.
OpenAI-compatible endpoint
The Gemini 3.1 Flash-Lite API speaks the OpenAI chat-completions wire format. Point any OpenAI SDK at RouterBase — no Google SDK or separate credentials needed.
One key for 200+ models
The same RouterBase key that calls the Gemini 3.1 Flash-Lite API also routes to GPT-5, Claude Opus 4.8, Gemini 3.1 Pro, and 200+ other models — no per-provider credential management.

Get started with the Gemini 3.1 Flash-Lite API in 3 steps
From sign-up to your first response in under 5 minutes.
Create a RouterBase API key
Sign up and generate an API key — one key reaches the Gemini 3.1 Flash-Lite API and every other model in the catalog.
Send your first message
Point any OpenAI-compatible SDK at routerbase.com/v1 and set the model to gemini-3-1-flash-lite. The response follows the standard chat-completions schema — streaming and multimodal input supported.
Inspect usage
Every response includes a detailed token breakdown — input, output, and cache-read tokens — so cost and cache-hit rate are visible on every call.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
What teams build with the Gemini 3.1 Flash-Lite API
Real production loads running on the Gemini 3.1 Flash-Lite API and the RouterBase model catalog.
Native multimodal on the Gemini 3.1 Flash-Lite API reads video and audio directly — one fast call transcribes and reasons over the content together.
Structured JSON outputs from the Gemini 3.1 Flash-Lite API killed our flaky regex parsing. Strictly valid JSON every time, zero retries.
We moved classification and routing to the Gemini 3.1 Flash-Lite API and latency dropped while quality climbed. RouterBase makes it 5% cheaper and routes around outages automatically.
1M tokens of context at Flash-Lite prices means I can throw whole docs at it cheaply. The Gemini 3.1 Flash-Lite API just handles them fast.
RouterBase puts the Gemini 3.1 Flash-Lite API and 200+ other models behind one key. Our team stopped filing requests for new provider accounts.
We A/B the Gemini 3.1 Flash-Lite API against 3.1 Pro per request — same SDK, one model field. Most traffic stays on Flash-Lite and the bill dropped sharply.
We routed our high-volume endpoints to the Gemini 3.1 Flash-Lite API over a weekend. The only PR comment was 'wait, that's all?'.
At $0.2375 / 1M input minus 5%, the Gemini 3.1 Flash-Lite API gave us newest-gen multimodal inference at a price that scales. The ROI math was instant.
Frequently Asked Questions
Common questions about the Gemini 3.1 Flash-Lite API.
It is RouterBase's pass-through to Google's Gemini 3.1 Flash-Lite — Google's fastest, most cost-effective Gemini 3.1 model, with a 1M-token context, native vision/audio/video input, function calling, and structured outputs, served via an OpenAI-compatible REST interface.