Chinese AI API Pricing
Prices normalized to USD per 1 million tokens, with currency, region, billing mode, effective dates and official sources preserved. Pricing is time-sensitive — always check the verification date.
| Provider | Model | Input $/1M | Output $/1M | Cached input $/1M | Batch $/1M | Effective | Note | Verified | Source |
|---|---|---|---|---|---|---|---|---|---|
| alibaba-cloud | qwen3.8-max | 2 | 6 | 0.25 | — | — | Beijing and Global regions (HK/Frankfurt/US/Tokyo): $1.65 input / $4.951 output. Beijing batch: 50% off. Explicit cache creation $2.5, explicit cache read $0.17 (Singapore). Free quota: 1M tokens, 90 days (Singapore). | 2026-09-20 | link |
| alibaba-cloud | qwen3.8-flash | 0.15 | 0.47 | 0.016 | — | — | Beijing and Global regions: $0.113 / $0.382. Explicit cache creation $0.2, read $0.016 (Singapore). Batch inference not supported. | 2026-09-20 | link |
| alibaba-cloud | qwen3.7-plus | 0.4 | 1.6 | — | — | — | Up to 256K context tier. 256K-1M tier: $1.2 / $4.8. Beijing: $0.276/$1.101 and $0.826/$3.301. Limited-time 20% console discount. Thinking billed same as output. | 2026-09-20 | link |
| bytedance | doubao-seed-2-1-pro | 6 | 30 | 1.2 | — | — | Standard tier. Low-priority tier ~50%: 3.00/15.00. Context cache storage 0.017 CNY per 1M tokens per hour. | 2026-09-20 | link |
| bytedance | doubao-seed-evolving | 6 | 30 | 1.2 | — | — | Standard tier. Pricing row matches Doubao Seed 2.1 Pro. | 2026-09-20 | link |
| bytedance | doubao-seed-2-1-turbo | 3 | 15 | 0.6 | — | — | Standard tier. Low-latency tier: 6.00/30.00. Low-priority tier: 1.50/7.50. | 2026-09-20 | link |
| deepseek | deepseek-v4-1-flash | 0.15 | 0.6 | 0.003 | — | — | Off-peak rates (all times except 01:00-04:00 and 06:00-10:00 UTC Mon-Fri). Peak = 2x: $0.30 input / $1.20 output / $0.006 cache hit. Concurrency limit 2500. No batch pricing listed. | 2026-09-20 | link |
| deepseek | deepseek-v4-pro | 0.66 | 1.98 | 0.022 | — | — | Off-peak rates; peak = 2x: $1.32 input / $3.96 output / $0.044 cache hit. Concurrency limit 500. Deprecation announced 2026-09-10 (service continuation per change log). | 2026-09-20 | link |
| minimax | minimax-m3 | 0.3 | 1.2 | 0.06 | — | — | Standard tier, <=512K input (permanent 50% off vs list $0.60/$2.40). >512K input: $0.60/$2.40, cache $0.12. Priority tier (service_tier=priority) = 1.5x. China platform: ¥2.1 / ¥8.4 (<=512K), ¥4.2 / ¥16.8 (>512K). | 2026-09-20 | link |
| minimax | minimax-m2.7 | 0.3 | 1.2 | 0.06 | — | — | Cache write $0.375 per 1M tokens. China platform: ¥2.1 / ¥8.4. | 2026-09-20 | link |
| minimax | minimax-m2.7-highspeed | 0.6 | 2.4 | 0.06 | — | — | Cache write $0.375 per 1M tokens. China platform: ¥4.2 / ¥16.8. | 2026-09-20 | link |
| zhipu-ai | glm-5.3 | 1.4 | 4.4 | 0.26 | — | — | Cache storage limited-time free. China platform (BigModel): ¥8 / ¥28, cached hit ¥2. Batch API = 50% of standard price for supported models. | 2026-09-20 | link |
| zhipu-ai | glm-5.3-flash | 0.15 | 0.5 | 0.03 | — | — | China platform: ¥0.8 / ¥2.8, cached ¥0.23. | 2026-09-20 | link |
| zhipu-ai | glm-5.3-flashx | 0.37 | 1.25 | 0.075 | — | — | China platform: ¥2 / ¥7, cached ¥0.57. Not yet on the GLM Coding Plan. | 2026-09-20 | link |
| zhipu-ai | glm-5.2 | 1.4 | 4.4 | 0.26 | — | — | China platform: ¥8 / ¥28, cached ¥2. | 2026-09-20 | link |
| moonshot-ai | kimi-k3 | 3 | 15 | 0.3 | — | — | Cache write: $3.00 (TTL 5min) or $6.00 (TTL 1h). Context 1,048,576. Access requires min $1 top-up. | 2026-09-20 | link |
| moonshot-ai | kimi-k2.7-code | 0.95 | 4 | 0.19 | — | — | Context 262,144. Thinking always on. | 2026-09-20 | link |
| moonshot-ai | kimi-k2.7-code-highspeed | 1.9 | 8 | 0.38 | — | — | Same model as k2.7-code at ~180 tokens/s. Context 262,144. | 2026-09-20 | link |
| moonshot-ai | kimi-k2.6 | 0.95 | 4 | 0.16 | — | — | Visual + text; thinking and non-thinking modes. Context 262,144. | 2026-09-20 | link |
All prices in original currency per 1M tokens unless noted. Always verify against the official pricing page before procurement decisions.