Benchmarks / Terminal-Bench
Terminal-Bench
Terminal-based agent benchmark (shell commands, file operations, package management and other command-line tasks).
Terminal-based agent benchmark (shell commands, file operations, package management and other command-line tasks). The table below lists 9 recorded evaluations across 9 models: deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, minimax-m3, glm-5.3, glm-5.2, kimi-k2.5, minimax-m2. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.
Benchmark methodology
| Task type | Terminal-based agent tasks (shell commands, file operations, compilation, package management, server setup) |
|---|---|
| Dataset size | ~100 tasks (beta release) |
| Evaluation method | Sandboxed terminal environment (Docker); each task has an English instruction, a test script verifying completion, and a reference (oracle) solution; agents run end-to-end autonomously |
| Scoring | Binary pass/fail per task; accuracy = share of tasks completed successfully |
Results
| Model | Model version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| deepseek-v4-1-flash | 2.1 | 90.6 | accuracy | 2026-09-10 | vendor_reported | link |
| deepseek-v4-pro | 2.1 | 87.9 | accuracy | 2026-08-13 | vendor_reported | link |
| qwen3.8-max | 2.1 | 86.6 | accuracy | 2026-08 | vendor_reported | link |
| kimi-k3 | 2.1 | 88.3 | accuracy | 2026-07 | vendor_reported | link |
| minimax-m3 | 2.1 | 66 | accuracy | 2026-06-01 | vendor_reported | link |
| glm-5.3 | 3.0 | 28.3 | accuracy | 2026-08-18 | vendor_reported | link |
| glm-5.2 | 3.0 | 4.6 | accuracy | 2026-08-18 | vendor_reported | link |
| kimi-k2.5 | 2.0 | 50.8 | accuracy | — | vendor_reported | link |
| minimax-m2 | version not stated in source | 46.3 | accuracy | 2025-10 | vendor_reported | link |
Limitations
Benchmark versions (2.0 / 2.1 / 3.0) are not comparable to each other; the version is recorded per evaluation. All scores are vendor-reported and not independently verified.
Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.
Sources
What does Terminal-Bench measure?
Terminal-based agent benchmark (shell commands, file operations, package management and other command-line tasks).
Which Chinese AI models have published Terminal-Bench results?
deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, minimax-m3, glm-5.3, glm-5.2, kimi-k2.5, minimax-m2.
Are Terminal-Bench scores independently verified?
Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.
Where does China AI Hub get its Terminal-Bench data?
From 4 sources, last verified 2026-09-20.