China AI Hub China AI Hub

Benchmarks / Terminal-Bench

Terminal-Bench

Terminal-based agent benchmark (shell commands, file operations, package management and other command-line tasks).

Terminal-Bench
Image: AI-generated illustration (Seedream)

Terminal-based agent benchmark (shell commands, file operations, package management and other command-line tasks). The table below lists 9 recorded evaluations across 9 models: deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, minimax-m3, glm-5.3, glm-5.2, kimi-k2.5, minimax-m2. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.

Benchmark methodology

Methodology of Terminal-Bench
Task type Terminal-based agent tasks (shell commands, file operations, compilation, package management, server setup)
Dataset size ~100 tasks (beta release)
Evaluation method Sandboxed terminal environment (Docker); each task has an English instruction, a test script verifying completion, and a reference (oracle) solution; agents run end-to-end autonomously
Scoring Binary pass/fail per task; accuracy = share of tasks completed successfully
Last verified: · Data status: Current · Next review:

Results

Results for Terminal-Bench
Model Model version Score Metric Date Source type Source
deepseek-v4-1-flash 2.1 90.6 accuracy 2026-09-10 vendor_reported link
deepseek-v4-pro 2.1 87.9 accuracy 2026-08-13 vendor_reported link
qwen3.8-max 2.1 86.6 accuracy 2026-08 vendor_reported link
kimi-k3 2.1 88.3 accuracy 2026-07 vendor_reported link
minimax-m3 2.1 66 accuracy 2026-06-01 vendor_reported link
glm-5.3 3.0 28.3 accuracy 2026-08-18 vendor_reported link
glm-5.2 3.0 4.6 accuracy 2026-08-18 vendor_reported link
kimi-k2.5 2.0 50.8 accuracy vendor_reported link
minimax-m2 version not stated in source 46.3 accuracy 2025-10 vendor_reported link

Limitations

Benchmark versions (2.0 / 2.1 / 3.0) are not comparable to each other; the version is recorded per evaluation. All scores are vendor-reported and not independently verified.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does Terminal-Bench measure?

Terminal-based agent benchmark (shell commands, file operations, package management and other command-line tasks).

Which Chinese AI models have published Terminal-Bench results?

deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, minimax-m3, glm-5.3, glm-5.2, kimi-k2.5, minimax-m2.

Are Terminal-Bench scores independently verified?

Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its Terminal-Bench data?

From 4 sources, last verified 2026-09-20.