China AI Hub China AI Hub

Benchmarks / SWE-bench

SWE-bench

Software engineering benchmark family built from real GitHub issues, with Pro, Verified and Multilingual variants.

SWE-bench
Image: AI-generated illustration (Seedream)

Software engineering benchmark family built from real GitHub issues, with Pro, Verified and Multilingual variants. The table below lists 5 recorded evaluations across 4 models: qwen3.8-max, minimax-m3, kimi-k2.5, minimax-m2. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.

Benchmark methodology

Methodology of SWE-bench
Task type Software engineering (resolve real GitHub issues by generating code patches)
Dataset size 2,294 task instances from 12 Python repositories (full set); SWE-bench Verified = 500 human-confirmed solvable problems
Evaluation method Docker-containerized harness; the model's patch is applied to the repository and the project's tests are run to verify resolution
Scoring Resolved rate (% of instances where all tests pass)
Last verified: · Data status: Current · Next review:

Results

Results for SWE-bench
Model Model version Score Metric Date Source type Source
qwen3.8-max Pro 67.7 accuracy 2026-08 vendor_reported link
minimax-m3 Pro 59 accuracy 2026-06-01 vendor_reported link
kimi-k2.5 Verified 76.8 accuracy vendor_reported link
minimax-m2 Verified 69.4 accuracy 2025-10 vendor_reported link
minimax-m2 Multilingual 56.5 accuracy 2025-10 vendor_reported link

Limitations

Pro, Verified and Multilingual variants are different test sets and are not comparable to each other; the variant is recorded per evaluation. All scores are vendor-reported and not independently verified.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does SWE-bench measure?

Software engineering benchmark family built from real GitHub issues, with Pro, Verified and Multilingual variants.

Which Chinese AI models have published SWE-bench results?

qwen3.8-max, minimax-m3, kimi-k2.5, minimax-m2.

Are SWE-bench scores independently verified?

Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its SWE-bench data?

From 4 sources, last verified 2026-09-20.