China AI Hub China AI Hub

Benchmarks / AutomationBench

AutomationBench

Benchmark of computer-use automation tasks.

AutomationBench
Image: AI-generated illustration (Seedream)

Benchmark of computer-use automation tasks. The table below lists 4 recorded evaluations across 4 models: deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, glm-5.3-flash. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.

Benchmark methodology

Methodology of AutomationBench
Task type End-to-end business workflow automation in simulated SaaS environments
Dataset size 47 simulated SaaS tools across 6 business functions (Sales, Marketing, Operations, Support, Finance, HR)
Evaluation method Each task initializes a simulated business environment (CRM, calendar, inbox); the agent must leave the environment in the correct end state
Scoring Task success rate (environment end-state verification)
Last verified: · Data status: Current · Next review:

Results

Results for AutomationBench
Model Model version Score Metric Date Source type Source
deepseek-v4-1-flash 54.8 accuracy 2026-09-10 vendor_reported link
deepseek-v4-pro Public 31.8 accuracy 2026-08-13 vendor_reported link
qwen3.8-max Pass@1 27.3 accuracy 2026-08 vendor_reported link
glm-5.3-flash GLM-5.2: 26.2 48.8 accuracy 2026-08-26 vendor_reported link

Limitations

All scores are vendor-reported and not independently verified. Pass@1 vs other sampling settings differ between vendors.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does AutomationBench measure?

Benchmark of computer-use automation tasks.

Which Chinese AI models have published AutomationBench results?

deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, glm-5.3-flash.

Are AutomationBench scores independently verified?

Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its AutomationBench data?

From 3 sources, last verified 2026-09-20.