Our support agents send GLM-4.6V a screenshot of what the user sees and it pinpoints the broken element — far cheaper than the vision model we ran before.
The GLM 4.6V API is Z.ai's 4.6-generation vision-language model — it takes images, video frames, and documents alongside text for OCR, chart and table extraction, and visual reasoning, with native tool calling and structured outputs. The GLM 4.6V API is OpenAI-compatible: point any existing SDK at RouterBase and set the model to glm-4-6v. Input at $0.255 / 1M tokens, output at $0.765 / 1M tokens, cached-input reads at $0.0425 / 1M — 15% below the official rate, all through one REST endpoint.
GLM 4.6V API: Chat
Use the GLM 4.6V API to add eyes to your app — Z.ai's 4.6-generation vision-language model reads images, video frames, screenshots, and documents, then answers in text or tool calls.
The GLM 4.6V API is Z.ai's flagship vision-language GLM model — a frontier large language model built for image understanding, visual reasoning, and multimodal chat, with native tool calling, structured outputs, and a long context window. Routed through RouterBase, the GLM 4.6V API is fully OpenAI-compatible: point any existing SDK at RouterBase's base URL, set the model to glm-4-6v, and your first GLM 4.6V API call goes through immediately.
Pricing for the GLM 4.6V API is $0.255 / 1M input tokens, $0.765 / 1M output tokens, and $0.0425 / 1M cached-input reads — 15% below the official published rate. One RouterBase key reaches the GLM 4.6V API and 200+ other models — no Z.ai account required.
Six reasons teams ship on the GLM 4.6V API
From frame-accurate image reading to a 15%-off price — why teams put GLM-4.6V behind their multimodal features.
Sees images, video & docs
The GLM 4.6V API takes images, multi-page documents, and video frames alongside text — handling OCR, chart and table extraction, and visual reasoning in one multimodal call.
Documents, charts & OCR
Point the GLM 4.6V API at a screenshot, scanned PDF, invoice, or dashboard and it pulls out text, tables, and structure — no separate OCR pipeline to run.
Native tool calling
The GLM 4.6V API calls functions in the OpenAI tool format straight from what it sees — let it read a form, then trigger your tools in one multimodal agent loop.
Long context window
Fit long multi-image documents, full chat history, and interleaved image-and-text turns into a single GLM 4.6V API request.
OpenAI-compatible, one key
The GLM 4.6V API speaks the OpenAI chat-completions format, images included — and the same RouterBase key also routes to GPT-5.5, Claude, Gemini, and 200+ other models.
15% off the list price
The GLM 4.6V API runs 15% under Z.ai's published rate — $0.255 / 1M input, $0.765 / 1M output, $0.0425 / 1M cached. Image input is billed as tokens, so multimodal stays cheap.

Get started with the GLM 4.6V API in 3 steps
From sign-up to your first response in under 5 minutes.
Create a RouterBase API key
Create an account and generate a key — it unlocks GLM-4.6V plus every other model in the RouterBase catalog.
Send your first message
Point any OpenAI-compatible SDK at routerbase.com/v1, set the model to glm-4-6v, and send messages with image_url parts next to text. Responses follow the standard chat-completions schema, streaming included.
Inspect usage
Each response carries a full token breakdown — text and image input, output, and cached-input — so you can watch multimodal cost on the GLM 4.6V API call by call.
Pay only for what you use
RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.
What teams build with the GLM 4.6V API
Real production loads running on the RouterBase model catalog across 200+ models.
We shipped receipt and invoice scanning in a weekend — one base-URL swap to the GLM 4.6V API and the image parsing just worked, no OCR vendor.
Our robots stream camera frames to GLM-4.6V and it returns grounded actions through standard tool calls — the OpenAI format meant zero adapter work.
We benchmarked GLM-4.6V against bigger vision models behind one RouterBase key — on chart and table reading it held its own at a fraction of the price.
Our RAG indexes scanned contracts as page images now — GLM-4.6V reads them directly, and the cached-input rate keeps long reference docs cheap.
Interleaving images and text in one message just works with the GLM 4.6V API — we send a chart plus a question and get a grounded answer back.
Frequently Asked Questions
Common questions about the GLM 4.6V API.
The GLM 4.6V API is RouterBase's pass-through to Z.ai's GLM-4.6V — a 4.6-generation vision-language model that takes images, video frames, and documents as input, served over an OpenAI-compatible REST interface with native tool calling and structured outputs.