chat.glm_4_6v
zai/glm-4-6v

The GLM 4.6V API is Z.ai's 4.6-generation vision-language model — it takes images, video frames, and documents alongside text for OCR, chart and table extraction, and visual reasoning, with native tool calling and structured outputs. The GLM 4.6V API is OpenAI-compatible: point any existing SDK at RouterBase and set the model to glm-4-6v. Input at $0.255 / 1M tokens, output at $0.765 / 1M tokens, cached-input reads at $0.0425 / 1M — 15% below the official rate, all through one REST endpoint.

Input
GLM-4.6V
max_tokens4096
temperature1
top_p1
presence_penalty0
frequency_penalty0
GLM-4.6V Online
Hi! I'm a helpful AI assistant. What can I do for you?

GLM 4.6V API: Chat

Use the GLM 4.6V API to add eyes to your app — Z.ai's 4.6-generation vision-language model reads images, video frames, screenshots, and documents, then answers in text or tool calls.

The GLM 4.6V API is Z.ai's flagship vision-language GLM model — a frontier large language model built for image understanding, visual reasoning, and multimodal chat, with native tool calling, structured outputs, and a long context window. Routed through RouterBase, the GLM 4.6V API is fully OpenAI-compatible: point any existing SDK at RouterBase's base URL, set the model to glm-4-6v, and your first GLM 4.6V API call goes through immediately.

Pricing for the GLM 4.6V API is $0.255 / 1M input tokens, $0.765 / 1M output tokens, and $0.0425 / 1M cached-input reads — 15% below the official published rate. One RouterBase key reaches the GLM 4.6V API and 200+ other models — no Z.ai account required.

Why this model

Six reasons teams ship on the GLM 4.6V API

From frame-accurate image reading to a 15%-off price — why teams put GLM-4.6V behind their multimodal features.

Sees images, video & docs

The GLM 4.6V API takes images, multi-page documents, and video frames alongside text — handling OCR, chart and table extraction, and visual reasoning in one multimodal call.

Documents, charts & OCR

Point the GLM 4.6V API at a screenshot, scanned PDF, invoice, or dashboard and it pulls out text, tables, and structure — no separate OCR pipeline to run.

Native tool calling

The GLM 4.6V API calls functions in the OpenAI tool format straight from what it sees — let it read a form, then trigger your tools in one multimodal agent loop.

Long context window

Fit long multi-image documents, full chat history, and interleaved image-and-text turns into a single GLM 4.6V API request.

OpenAI-compatible, one key

The GLM 4.6V API speaks the OpenAI chat-completions format, images included — and the same RouterBase key also routes to GPT-5.5, Claude, Gemini, and 200+ other models.

15% off the list price

The GLM 4.6V API runs 15% under Z.ai's published rate — $0.255 / 1M input, $0.765 / 1M output, $0.0425 / 1M cached. Image input is billed as tokens, so multimodal stays cheap.

RouterBase dashboard preview
Quickstart

Get started with the GLM 4.6V API in 3 steps

From sign-up to your first response in under 5 minutes.

  1. Create a RouterBase API key

    Create an account and generate a key — it unlocks GLM-4.6V plus every other model in the RouterBase catalog.

  2. Send your first message

    Point any OpenAI-compatible SDK at routerbase.com/v1, set the model to glm-4-6v, and send messages with image_url parts next to text. Responses follow the standard chat-completions schema, streaming included.

  3. Inspect usage

    Each response carries a full token breakdown — text and image input, output, and cached-input — so you can watch multimodal cost on the GLM 4.6V API call by call.

Pricing

Pay only for what you use

RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.

Customer stories

What teams build with the GLM 4.6V API

Real production loads running on the RouterBase model catalog across 200+ models.

Marcus Reyes
Marcus ReyesCTO, Paradigm AI

Our support agents send GLM-4.6V a screenshot of what the user sees and it pinpoints the broken element — far cheaper than the vision model we ran before.

Priya Lakshmi
Priya LakshmiFounder, Quillo

We shipped receipt and invoice scanning in a weekend — one base-URL swap to the GLM 4.6V API and the image parsing just worked, no OCR vendor.

Aoi Tanaka
Aoi TanakaML Lead, Daybreak Robotics

Our robots stream camera frames to GLM-4.6V and it returns grounded actions through standard tool calls — the OpenAI format meant zero adapter work.

Ethan Nguyen
Ethan NguyenHead of Engineering, Compound Studio

We benchmarked GLM-4.6V against bigger vision models behind one RouterBase key — on chart and table reading it held its own at a fraction of the price.

Sophia Martín
Sophia MartínCTO, Relay

Our RAG indexes scanned contracts as page images now — GLM-4.6V reads them directly, and the cached-input rate keeps long reference docs cheap.

Lucas Fernandes
Lucas FernandesEngineering Manager, Light

Interleaving images and text in one message just works with the GLM 4.6V API — we send a chart plus a question and get a grounded answer back.

Frequently Asked Questions

Common questions about the GLM 4.6V API.

The GLM 4.6V API is RouterBase's pass-through to Z.ai's GLM-4.6V — a 4.6-generation vision-language model that takes images, video frames, and documents as input, served over an OpenAI-compatible REST interface with native tool calling and structured outputs.