chat.gemini_3_flash_preview
google/gemini-3-flash-preview

The Gemini 3 Flash (Preview) API is Google's fast, low-cost multimodal model — a 1M-token context window, native input for text, images, audio, video and PDF, reasoning, function calling, and structured JSON outputs. The Gemini 3 Flash Preview API is OpenAI-compatible: point any existing SDK at RouterBase and set the model to gemini-3-flash-preview. Input at $0.475 / 1M tokens, output at $2.85 / 1M tokens, cache reads at $0.0475 / 1M tokens — all through one REST endpoint.

Input
Gemini 3 Flash (Preview)
max_tokens4096
temperature1
top_p1
presence_penalty0
frequency_penalty0
Gemini 3 Flash (Preview) Online
Hi! I'm a helpful AI assistant. What can I do for you?

Gemini 3 Flash Preview API: Chat

Use the Gemini 3 Flash Preview API to run Google's fast, low-cost Gemini 3 multimodal model — a 1M-token context, native vision/audio/video input, and structured outputs.

The Gemini 3 Flash Preview API is Google's fast, cost-efficient Gemini 3 multimodal model — a 1M-token context window, native input for text, images, audio, video, and PDF, function calling, and structured JSON outputs, tuned for low latency and high throughput. Routed through RouterBase, the Gemini 3 Flash Preview API is fully OpenAI-compatible: point any existing SDK at RouterBase's base URL, set the model to gemini-3-flash-preview, and your first Gemini 3 Flash Preview API call goes through immediately.

Pricing for the Gemini 3 Flash Preview API is $0.475 / 1M input tokens, $2.85 / 1M output tokens, and $0.0475 / 1M cached-input reads — 5% below the official published rate. One RouterBase key — no Google credentials required.

Why this model

Six reasons teams ship on the Gemini 3 Flash Preview API

From newest-gen speed to 5%-off pricing — what makes the Gemini 3 Flash Preview API stand out.

Fast, balanced Gemini 3

The Gemini 3 Flash Preview API is Google’s newest-generation workhorse — tuned for the best blend of speed, quality, and cost in the Gemini 3 family.

1M-token context

Accepts up to 1 million tokens per request — long documents, multi-file code, or video transcripts fit in a single Gemini 3 Flash Preview API call without chunking.

Native multimodal input

The Gemini 3 Flash Preview API understands text, images, audio, video, and PDFs. Pass mixed media in one request to analyze, transcribe, or reason across modalities.

Function calling & JSON

Define tools and request a JSON schema; the Gemini 3 Flash Preview API returns structured tool calls and strictly valid JSON for reliable agent and extraction pipelines.

OpenAI-compatible endpoint

The Gemini 3 Flash Preview API speaks the OpenAI chat-completions wire format. Point any OpenAI SDK at RouterBase — no Google SDK or separate credentials needed.

One key for 200+ models

The same RouterBase key that calls the Gemini 3 Flash Preview API also routes to GPT-5, Claude Opus 4.8, Gemini 3 Pro, and 200+ other models — no per-provider credential management.

RouterBase dashboard preview
Quickstart

Get started with the Gemini 3 Flash Preview API in 3 steps

From sign-up to your first response in under 5 minutes.

  1. Create a RouterBase API key

    Sign up and generate an API key — one key reaches the Gemini 3 Flash Preview API and every other model in the catalog.

  2. Send your first message

    Point any OpenAI-compatible SDK at routerbase.com/v1 and set the model to gemini-3-flash-preview. The Gemini 3 Flash Preview API response follows the standard chat-completions schema — streaming and multimodal input supported.

  3. Inspect usage

    Every Gemini 3 Flash Preview API response includes a detailed token breakdown — input, output, and cache-read tokens — so cost and cache-hit rate are visible on every call.

Pricing

Pay only for what you use

RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.

Customer stories

What teams build with the Gemini 3 Flash Preview API

Real production loads running on the Gemini 3 Flash Preview API and the RouterBase model catalog.

Marcus Reyes
Marcus ReyesCTO, Paradigm AI

The Gemini 3 Flash Preview API is our everyday default — newest-gen quality, fast, and cheap enough to skip the flagship on most calls. The 5% RouterBase discount is pure margin.

Priya Lakshmi
Priya LakshmiFounder, Quillo

Native multimodal on the Gemini 3 Flash Preview API reads video and audio directly — one fast call transcribes and reasons over the content together.

Thomas Beck
Thomas BeckStaff Engineer, Northbeam

Structured JSON outputs from the Gemini 3 Flash Preview API killed our flaky regex parsing. Strictly valid JSON every time, zero retries.

Aoi Tanaka
Aoi TanakaML Lead, Daybreak Robotics

We moved classification and routing to the Gemini 3 Flash Preview API and latency dropped while quality climbed. RouterBase makes it 5% cheaper and routes around outages automatically.

Jonas Keller
Jonas KellerIndie Developer

1M tokens of context at Flash prices means I can throw whole docs at it cheaply. The Gemini 3 Flash Preview API just handles them fast.

Ethan Nguyen
Ethan NguyenHead of Engineering, Compound Studio

RouterBase puts the Gemini 3 Flash Preview API and 200+ other models behind one key. Our team stopped filing requests for new provider accounts.

Sophia Martín
Sophia MartínCTO, Relay

We A/B the Gemini 3 Flash Preview API against 3 Pro per request — same SDK, one model field. Most traffic stays on Flash and the bill dropped sharply.

Lucas Fernandes
Lucas FernandesEngineering Manager, Light

We routed our high-volume endpoints to the Gemini 3 Flash Preview API over a weekend. The only PR comment was 'wait, that's all?'.

David Okonkwo
David OkonkwoCo-founder, Figment

At $0.475 / 1M input minus 5%, the Gemini 3 Flash Preview API gave us newest-gen multimodal inference at a price that scales. The ROI math was instant.

Frequently Asked Questions

Common questions about the Gemini 3 Flash Preview API.

It is RouterBase's pass-through to Google's Gemini 3 Flash (Preview) — Google's fast, cost-efficient Gemini 3 model, with a 1M-token context, native vision/audio/video input, function calling, and structured outputs, served via an OpenAI-compatible REST interface.