China AI Hub China AI Hub

Benchmarks / DeepSWE

DeepSWE

Software engineering benchmark built from real-world issues and pull requests.

DeepSWE
Image: AI-generated illustration (Seedream)

Software engineering benchmark built from real-world issues and pull requests. The table below lists 6 recorded evaluations across 6 models: deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, glm-5.3, glm-5.3-flash. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.

Benchmark methodology

Methodology of DeepSWE
Task type Long-horizon software engineering tasks on active open-source repositories
Dataset size 113 tasks across TypeScript, Go, Python, JavaScript and Rust
Evaluation method Isolated agent environment with program-based verifiers; the agent's committed patch is applied and graded in a pristine container
Scoring Binary reward plus pass fractions per task (reward.json / CTRF test report)
Last verified: · Data status: Current · Next review:

Results

Results for DeepSWE
Model Model version Score Metric Date Source type Source
deepseek-v4-1-flash v1.1 74.2 accuracy 2026-09-10 vendor_reported link
deepseek-v4-pro version not stated in source 62.7 accuracy 2026-08-13 vendor_reported link
qwen3.8-max v1.1 56.6 accuracy 2026-08 vendor_reported link
kimi-k3 67.3 with mini-SWE-agent harness 67.5 accuracy 2026-07 vendor_reported link
glm-5.3 v1.1 66.9 accuracy 2026-08-18 vendor_reported link
glm-5.3-flash v1.1 63.4 accuracy 2026-08-26 vendor_reported link

Limitations

All scores are vendor-reported and not independently verified. Some vendors do not state the benchmark version; unversioned scores should not be compared with versioned ones.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does DeepSWE measure?

Software engineering benchmark built from real-world issues and pull requests.

Which Chinese AI models have published DeepSWE results?

deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, glm-5.3, glm-5.3-flash.

Are DeepSWE scores independently verified?

Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its DeepSWE data?

From 4 sources, last verified 2026-09-20.