Benchmarks / SWE-bench
SWE-bench
Software engineering benchmark family built from real GitHub issues, with Pro, Verified and Multilingual variants.
Software engineering benchmark family built from real GitHub issues, with Pro, Verified and Multilingual variants. The table below lists 5 recorded evaluations across 4 models: qwen3.8-max, minimax-m3, kimi-k2.5, minimax-m2. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.
Benchmark methodology
| Task type | Software engineering (resolve real GitHub issues by generating code patches) |
|---|---|
| Dataset size | 2,294 task instances from 12 Python repositories (full set); SWE-bench Verified = 500 human-confirmed solvable problems |
| Evaluation method | Docker-containerized harness; the model's patch is applied to the repository and the project's tests are run to verify resolution |
| Scoring | Resolved rate (% of instances where all tests pass) |
Results
| Model | Model version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| qwen3.8-max | Pro | 67.7 | accuracy | 2026-08 | vendor_reported | link |
| minimax-m3 | Pro | 59 | accuracy | 2026-06-01 | vendor_reported | link |
| kimi-k2.5 | Verified | 76.8 | accuracy | — | vendor_reported | link |
| minimax-m2 | Verified | 69.4 | accuracy | 2025-10 | vendor_reported | link |
| minimax-m2 | Multilingual | 56.5 | accuracy | 2025-10 | vendor_reported | link |
Limitations
Pro, Verified and Multilingual variants are different test sets and are not comparable to each other; the variant is recorded per evaluation. All scores are vendor-reported and not independently verified.
Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.
Sources
What does SWE-bench measure?
Software engineering benchmark family built from real GitHub issues, with Pro, Verified and Multilingual variants.
Which Chinese AI models have published SWE-bench results?
qwen3.8-max, minimax-m3, kimi-k2.5, minimax-m2.
Are SWE-bench scores independently verified?
Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.
Where does China AI Hub get its SWE-bench data?
From 4 sources, last verified 2026-09-20.