chat.qwen3_5_122b_a10b
alibaba/qwen3-5-122b-a10b

Qwen3.5 122B A10B is a Mixture-of-Experts LLM from Alibaba Qwen with 122B total and about 10B active parameters, a 256K context, and strong reasoning, coding, and multilingual skills. Served on RouterBase over an OpenAI-compatible endpoint at $0.34 / 1M input and $2.72 / 1M output, 15% below list.

Input
Qwen3.5 122B A10B
max_tokens4096
temperature1
top_p1
presence_penalty0
frequency_penalty0
Qwen3.5 122B A10B Online
Hi! I'm a helpful AI assistant. What can I do for you?

Qwen3.5 122B A10B API - 122B MoE, 10B Active, 256K Context

The Qwen3.5 122B A10B API runs a 122B MoE with about 10B active parameters, a 256K context, and strong reasoning, coding, and multilingual chat.

The Qwen3.5 122B A10B API serves Alibaba Qwen3.5 122B A10B, a Mixture-of-Experts model with 122 billion total parameters and roughly 10 billion active per token. Sparse routing gives you big-model quality at the latency and per-token cost of a much smaller network, and a 256K token context window fits long files, transcripts, and tool traces in one call. You reach the Qwen3.5 122B A10B API through RouterBase over the OpenAI chat-completions protocol, so your existing client works unchanged and one key covers the whole catalog.

Qwen3.5 122B A10B was tuned for multi-step reasoning, code generation and review, agentic tool use, structured JSON, and multilingual chat across 100+ languages. Billing is $0.34 per 1M input tokens and $2.72 per 1M output tokens, 15% below the list rate of $0.40 in and $3.20 out, with no per-request fee. Streaming and standard OpenAI tool calls mean the Qwen3.5 122B A10B API drops into agent loops, IDE assistants, and batch pipelines without provider-specific glue.

Sparse, long-context, agent-ready

What the Qwen3.5 122B A10B API gives you

Frontier-class quality on a 10B active budget.

122B MoE, 10B active

The Qwen3.5 122B A10B API routes each token through a small set of experts inside a 122B-parameter network, so quality lands near much larger dense models while latency stays low.

256K context

A 256K token window lets the model hold long files, transcripts, and tool traces in a single call, so you rarely have to chunk inputs to the Qwen3.5 122B A10B API.

Reasoning that holds up

The Qwen3.5 generation sharpens multi-step reasoning and instruction following, so the Qwen3.5 122B A10B API works through hard problems without drifting off task.

Coding and agents

Strong code generation, review, and agentic tool use make the Qwen3.5 122B A10B API a dependable backbone for developer tools and autonomous workflows.

Tools and JSON

The Qwen3.5 122B A10B API returns function calls and structured JSON in the standard OpenAI schema, ready for agents and data pipelines.

Streaming, per token

The Qwen3.5 122B A10B API streams tokens as they generate and bills per token at $0.34 in and $2.72 out, 15% below list, with no request fee.

RouterBase dashboard preview
Get going

First call to the Qwen3.5 122B A10B API in three steps

Wire it once, then stream.

  1. Create a key

    One RouterBase key reaches the Qwen3.5 122B A10B API and every other model in the catalog, so there is nothing provider-specific to set up.

  2. Point your client

    Send any OpenAI client to routerbase.com/v1 with model=qwen/qwen3.5-122b-a10b and a messages array; the Qwen3.5 122B A10B API answers on the same chat-completions shape.

  3. Stream and watch usage

    The model streams tokens as they generate and returns token usage per response, so cost and latency stay visible on every call to the Qwen3.5 122B A10B API.

Pricing

Pay only for what you use

RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.

Pricing

Pay only for what you use

RouterBase passes through partner-tier pricing. Compared against the model's official published API rate.

ResolutionRouterBaseOfficialSave
Input (per 1M) $0.340$0.400−15%
Output (per 1M) $2.720$3.200−15%

One run at the default settings (Input (per 1M)) costs $0.340.With $1 you can run this model approximately 2 times.

Official price = Alibaba Qwen list rate for Qwen3.5 122B A10B ($0.40 per 1M input, $3.20 per 1M output). RouterBase serves the Qwen3.5 122B A10B API at $0.34 / $2.72, 15% below list.

Shipping on Qwen3.5

Teams building on the Qwen3.5 122B A10B API

Reasoning, agents, and chat behind one call.

Ingrid SolheimHead of AI, Fjordline

We moved our agent stack to the Qwen3.5 122B A10B API and the 10B active budget cut our p95 latency almost in half at the same quality.

Rafael OrtegaStaff Engineer, Lumen Works

The 256K context meant the Qwen3.5 122B A10B API read our whole design brief plus its tool traces in one pass - no more chunking.

Amara DialloCTO, Kite Analytics

Function calling in the standard schema meant our agents ran on the Qwen3.5 122B A10B API with zero custom parsing on day one.

Jonas Keller
Jonas KellerFounder, Brightpath

At $0.34 in and $2.72 out the Qwen3.5 122B A10B API kept our per-task cost predictable as traffic climbed through the quarter.

Mei-Ling ChouBackend Lead, Harborlight

Streaming from the Qwen3.5 122B A10B API is smooth, so our chat UI renders token by token without any extra buffering work.

Viktor AntonovML Lead, Northgate

We push code review and JSON extraction to the Qwen3.5 122B A10B API and the structured output comes back clean every time.

Sofia MarchettiPrincipal Engineer, Vellum Labs

Reasoning quality surprised us - the Qwen3.5 122B A10B API works through multi-step tickets that used to need a much pricier model.

Andre BoucherFounder, Quai Digital

Switching from another provider to the Qwen3.5 122B A10B API was one string change - the OpenAI shape did not move at all.

Lucia FerreiraEngineering Manager, Mirante

Multilingual coverage is why we standardized on the Qwen3.5 122B A10B API; Portuguese and English answers land equally well.

Qwen3.5 122B A10B API - common questions

What to know before you route traffic to it.

The Qwen3.5 122B A10B API is RouterBase hosted access to Alibaba Qwen3.5 122B A10B, a Mixture-of-Experts LLM with 122B total and 10B active parameters, over an OpenAI-compatible endpoint.