Models / GLM-5.3-Flash
GLM-5.3-Flash
GLM-5.3-Flash family · Provider: zhipu-ai · Status: active · Released: 2026-08-26
GLM-5.3-Flash is a GLM-5.3-Flash model developed by zhipu-ai, a Chinese AI company headquartered in Beijing Zhipu Huazhang Technology Co., Ltd. (北京智谱华章科技股份有限公司), Beijing, China, released 2026-08-26 with a 1,048,576-token context window and open weights.
| Key fact | Value |
|---|---|
| Model ID | glm-5.3-flash |
| Version | FlashX |
| Architecture | 320B total / 18B active; first open-source frontier model combining sparse + linear attention; mHC hyper-connections; 30T-token multimodal pre-training corpus |
| Context window | 1,048,576 tokens |
| Max output | 131,072 tokens |
| Open weights | Yes |
| License | Apache-2.0 (per GitHub repo metadata; README has no separate weights-license section - verify per-model HF cards before reuse) |
| Self-hosting | Yes |
| API available | Yes |
| API pricing | $0.15 input / $0.5 output per 1M tokens (USD) · provider pricing page |
| Regions | international, china |
| Cloud providers | Z.ai, BigModel |
Capabilities
| Capability | Supported |
|---|---|
| Reasoning | Yes |
| Coding | Yes |
| Math | Unknown |
| Chinese | Unknown |
| English | Unknown |
| Multilingual | Unknown |
| Vision | Yes |
| Audio | Unknown |
| Video | Yes |
| Tool calling | Unknown |
| Function calling | Unknown |
| Structured output | Unknown |
| Agent capability | Yes |
| RAG | Unknown |
| Computer use | Yes |
Benchmark results
| Benchmark | Version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| Artificial Analysis Intelligence Index v4.1.1 | — | 57 | index score | 2026-08-26 | vendor_reported | link |
| DeepSWE v1.1 | — | 63.4 | accuracy | 2026-08-26 | vendor_reported | link |
| AutomationBench | — | 48.8 | accuracy | 2026-08-26 | vendor_reported | link |
Benchmark scores are single data points, not universal rankings. Vendor-reported scores are labeled as such.
Known limitations
- FlashX tier not yet available on the GLM Coding Plan (pay-as-you-go only)
- Reasoning always enabled; cannot be disabled
- Z.ai Code Bench is a private in-house benchmark
- Benchmarks vendor-reported; not independently verified
GLM-5.3-Flash (2026-08-26) is Zhipu’s multimodal coding model with open weights (320B total / 18B active, FP8): 1M context, 128K max output, input modalities of video, image, text and file. It combines sparse and linear attention and is positioned for visual coding loops (observe - code - test), computer use (BUA/CUA), browser/GUI agents, office workflows, video understanding, 3D and CAD tasks. The FlashX tier serves at up to 200 tokens/s.
International pricing: Flash $0.15 input / $0.50 output per 1M tokens (cached $0.03); FlashX $0.37 / $1.25 (cached $0.075), as of 2026-09-20. Zhipu states all Flash traffic is served on Chinese AI chips.
Provider
Pricing
Benchmarks with results for this model
Agents built on this model
API
Data interpretation
| Field | Value | Evidence type |
|---|---|---|
| Context window | 1,048,576 tokens | Official |
| Architecture | 320B total / 18B active; first open-source frontier model combining sparse + linear attention; mHC hyper-connections; 30T-token multimodal pre-training corpus | Official |
| Open weights | Yes | Official |
| License | Apache-2.0 (per GitHub repo metadata; README has no separate weights-license section - verify per-model HF cards before reuse) | Official |
| API pricing | $0.15 / $0.5 per 1M tokens (USD) | Official |
| Artificial Analysis Intelligence Index v4.1.1 | 57 index score | Vendor-reported |
| DeepSWE v1.1 | 63.4 accuracy | Vendor-reported |
| AutomationBench | 48.8 accuracy | Vendor-reported |
| Family position | 1 model in the GLM-5.3-Flash family | China AI Hub analysis |
Evidence types: Official = vendor documentation, pricing pages or model cards. Vendor-reported = benchmark scores published by the vendor. China AI Hub analysis = derived from the database itself. See the sourcing policy.
Sources
Confidence and source hierarchy per the sourcing policy. Facts change; check the source before relying on this page.
What is the context window of GLM-5.3-Flash?
GLM-5.3-Flash has a 1,048,576-token context window and a maximum output of 131,072 tokens.
Is GLM-5.3-Flash open weight?
Yes — GLM-5.3-Flash weights are openly available under the Apache-2.0 (per GitHub repo metadata; README has no separate weights-license section - verify per-model HF cards before reuse) license.
How much does GLM-5.3-Flash cost through the API?
$0.15 per 1M input tokens and $0.5 per 1M output tokens (USD).
Where does China AI Hub get its GLM-5.3-Flash data?
From 3 sources (official pages first), last verified 2026-09-20. Benchmark scores are labeled by source type; see the sourcing policy for details.