Benchmarks / DeepSWE
DeepSWE
Software engineering benchmark built from real-world issues and pull requests.
Software engineering benchmark built from real-world issues and pull requests. The table below lists 6 recorded evaluations across 6 models: deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, glm-5.3, glm-5.3-flash. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.
Benchmark methodology
| Task type | Long-horizon software engineering tasks on active open-source repositories |
|---|---|
| Dataset size | 113 tasks across TypeScript, Go, Python, JavaScript and Rust |
| Evaluation method | Isolated agent environment with program-based verifiers; the agent's committed patch is applied and graded in a pristine container |
| Scoring | Binary reward plus pass fractions per task (reward.json / CTRF test report) |
Results
| Model | Model version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| deepseek-v4-1-flash | v1.1 | 74.2 | accuracy | 2026-09-10 | vendor_reported | link |
| deepseek-v4-pro | version not stated in source | 62.7 | accuracy | 2026-08-13 | vendor_reported | link |
| qwen3.8-max | v1.1 | 56.6 | accuracy | 2026-08 | vendor_reported | link |
| kimi-k3 | 67.3 with mini-SWE-agent harness | 67.5 | accuracy | 2026-07 | vendor_reported | link |
| glm-5.3 | v1.1 | 66.9 | accuracy | 2026-08-18 | vendor_reported | link |
| glm-5.3-flash | v1.1 | 63.4 | accuracy | 2026-08-26 | vendor_reported | link |
Limitations
All scores are vendor-reported and not independently verified. Some vendors do not state the benchmark version; unversioned scores should not be compared with versioned ones.
Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.
Sources
What does DeepSWE measure?
Software engineering benchmark built from real-world issues and pull requests.
Which Chinese AI models have published DeepSWE results?
deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, glm-5.3, glm-5.3-flash.
Are DeepSWE scores independently verified?
Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.
Where does China AI Hub get its DeepSWE data?
From 4 sources, last verified 2026-09-20.