China AI Hub China AI Hub

Models / GLM-5.3-Flash

GLM-5.3-Flash

GLM-5.3-Flash family · Provider: zhipu-ai · Status: active · Released: 2026-08-26

GLM-5.3-Flash is a GLM-5.3-Flash model developed by zhipu-ai, a Chinese AI company headquartered in Beijing Zhipu Huazhang Technology Co., Ltd. (北京智谱华章科技股份有限公司), Beijing, China, released 2026-08-26 with a 1,048,576-token context window and open weights.

GLM-5.3-Flash
Image: AI-generated illustration (Seedream)
Last verified: · Data status: Current · Next review:
Key facts for GLM-5.3-Flash
Key factValue
Model IDglm-5.3-flash
VersionFlashX
Architecture320B total / 18B active; first open-source frontier model combining sparse + linear attention; mHC hyper-connections; 30T-token multimodal pre-training corpus
Context window 1,048,576 tokens
Max output131,072 tokens
Open weightsYes
LicenseApache-2.0 (per GitHub repo metadata; README has no separate weights-license section - verify per-model HF cards before reuse)
Self-hostingYes
API availableYes
API pricing $0.15 input / $0.5 output per 1M tokens (USD) · provider pricing page
Regionsinternational, china
Cloud providersZ.ai, BigModel

Capabilities

Capabilities of GLM-5.3-Flash
CapabilitySupported
ReasoningYes
CodingYes
MathUnknown
ChineseUnknown
EnglishUnknown
MultilingualUnknown
VisionYes
AudioUnknown
VideoYes
Tool callingUnknown
Function callingUnknown
Structured outputUnknown
Agent capabilityYes
RAGUnknown
Computer useYes

Benchmark results

Benchmark results for GLM-5.3-Flash
BenchmarkVersionScoreMetricDateSource typeSource
Artificial Analysis Intelligence Index v4.1.1 57 index score 2026-08-26 vendor_reported link
DeepSWE v1.1 63.4 accuracy 2026-08-26 vendor_reported link
AutomationBench 48.8 accuracy 2026-08-26 vendor_reported link

Benchmark scores are single data points, not universal rankings. Vendor-reported scores are labeled as such.

Known limitations

GLM-5.3-Flash (2026-08-26) is Zhipu’s multimodal coding model with open weights (320B total / 18B active, FP8): 1M context, 128K max output, input modalities of video, image, text and file. It combines sparse and linear attention and is positioned for visual coding loops (observe - code - test), computer use (BUA/CUA), browser/GUI agents, office workflows, video understanding, 3D and CAD tasks. The FlashX tier serves at up to 200 tokens/s.

International pricing: Flash $0.15 input / $0.50 output per 1M tokens (cached $0.03); FlashX $0.37 / $1.25 (cached $0.075), as of 2026-09-20. Zhipu states all Flash traffic is served on Chinese AI chips.

Provider

Pricing

Benchmarks with results for this model

Agents built on this model

API

Data interpretation

Evidence types for GLM-5.3-Flash
FieldValueEvidence type
Context window1,048,576 tokensOfficial
Architecture320B total / 18B active; first open-source frontier model combining sparse + linear attention; mHC hyper-connections; 30T-token multimodal pre-training corpusOfficial
Open weightsYesOfficial
LicenseApache-2.0 (per GitHub repo metadata; README has no separate weights-license section - verify per-model HF cards before reuse)Official
API pricing $0.15 / $0.5 per 1M tokens (USD) Official
Artificial Analysis Intelligence Index v4.1.1 57 index score Vendor-reported
DeepSWE v1.1 63.4 accuracy Vendor-reported
AutomationBench 48.8 accuracy Vendor-reported
Family position1 model in the GLM-5.3-Flash familyChina AI Hub analysis

Evidence types: Official = vendor documentation, pricing pages or model cards. Vendor-reported = benchmark scores published by the vendor. China AI Hub analysis = derived from the database itself. See the sourcing policy.

Sources

Confidence and source hierarchy per the sourcing policy. Facts change; check the source before relying on this page.

What is the context window of GLM-5.3-Flash?

GLM-5.3-Flash has a 1,048,576-token context window and a maximum output of 131,072 tokens.

Is GLM-5.3-Flash open weight?

Yes — GLM-5.3-Flash weights are openly available under the Apache-2.0 (per GitHub repo metadata; README has no separate weights-license section - verify per-model HF cards before reuse) license.

How much does GLM-5.3-Flash cost through the API?

$0.15 per 1M input tokens and $0.5 per 1M output tokens (USD).

Where does China AI Hub get its GLM-5.3-Flash data?

From 3 sources (official pages first), last verified 2026-09-20. Benchmark scores are labeled by source type; see the sourcing policy for details.