Core capabilities
- 235B total / 22B active MoE with 128 experts
- 128K context and hybrid thinking control
- Turn-level /think and /no_think switching
- Open weights under Apache 2.0 with tool integration
Alibaba Qwen / open
The flagship open-weight Qwen3 MoE model with thinking/non-thinking modes and broad multilingual support.
Strong across Chinese, multilingual, coding and agent tasks for localized or self-hosted products in Asian markets. Qwen3 235B-A22B is the flagship open MoE with 235B total and 22B active parameters, 128 experts and 128K context. It supports hybrid thinking plus multilingual, coding, math and agent workflows.
Use SGLang or vLLM with the recommended reasoning parser, or a managed Model Studio endpoint. Keep prompt templates versioned, test non-thinking latency separately and validate MCP/tool outputs before execution.
Total and active parameters differ substantially; size hardware around the actual quantization and parallelism plan. Thinking content and tool calls require the correct chat template and parser. Quantized checkpoints and community runtimes can differ in context behavior and reasoning format.
Specifications and positioning checked against provider documentation.
Open official source↗