China AI Hub China AI Hub

Benchmark Database

Benchmark results with evaluation method, model version, source type (vendor-reported vs independent) and date. Scores are data points — never universal rankings.

AutomationBench

4 results · verified 2026-09-20

Benchmark of computer-use automation tasks.…

BrowseComp

3 results · verified 2026-09-20

Benchmark of browsing and retrieval ability: locating obscure information using web search and browsing.…

CyberGym

3 results · verified 2026-09-20

Cybersecurity agent benchmark focused on vulnerability discovery tasks.…

DeepSWE

6 results · verified 2026-09-20

Software engineering benchmark built from real-world issues and pull requests.…

GPQA Diamond

5 results · verified 2026-09-20

Graduate-level science question-answering benchmark (expert-level questions in biology, physics, chemistry).…

MMMU-Pro

2 results · verified 2026-09-20

Multimodal, multi-discipline understanding benchmark with college-level questions requiring reasoning.…

HLE

6 results · verified 2026-09-20

Humanity's Last Exam - a frontier benchmark of expert-level questions across disciplines, often reported with and without tool access.…

SWE-bench

5 results · verified 2026-09-20

Software engineering benchmark family built from real GitHub issues, with Pro, Verified and Multilingual variants.…

Terminal-Bench

9 results · verified 2026-09-20

Terminal-based agent benchmark (shell commands, file operations, package management and other command-line tasks).…

Video-MME

2 results · verified 2026-09-20

Video understanding benchmark spanning various video durations and domains.…