China AI Hub China AI Hub

Benchmarks / HLE

HLE

Humanity's Last Exam - a frontier benchmark of expert-level questions across disciplines, often reported with and without tool access.

HLE
Image: AI-generated illustration (Seedream)

Humanity’s Last Exam - a frontier benchmark of expert-level questions across disciplines, often reported with and without tool access. The table below lists 6 recorded evaluations across 6 models: deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, kimi-k2.5, minimax-m2. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.

Benchmark methodology

Methodology of HLE
Task type Expert-level academic Q&A across mathematics, humanities and natural sciences
Dataset size 2,500 questions across dozens of subjects
Evaluation method Multiple-choice and short-answer questions developed by subject-matter experts; multimodal; suitable for automated grading
Scoring Accuracy (% correct); frequently reported with and without tool access
Contamination notes Dataset includes a canary string (hle:3r2s:26b5c67b-...) to aid model builders in filtering the dataset from future training.
Last verified: · Data status: Current · Next review:

Results

Results for HLE
Model Model version Score Metric Date Source type Source
deepseek-v4-1-flash 39.1 on pure-text subset 36.8 accuracy 2026-09-10 vendor_reported link
deepseek-v4-pro 42.7 (60.0 with tools) accuracy 2026-08-13 vendor_reported link
qwen3.8-max 43.6 (56.2 with tools) accuracy 2026-08 vendor_reported link
kimi-k3 HLE-Full 43.5 (56.0 with tools) accuracy 2026-07 vendor_reported link
kimi-k2.5 HLE-Full 30.1 (50.2 with tools) accuracy vendor_reported link
minimax-m2 12.5 without tools / 31.8 with tools accuracy 2025-10 vendor_reported link

Limitations

All scores are vendor-reported and not independently verified. With-tools and without-tools results are not directly comparable; the setting is recorded per score.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does HLE measure?

Humanity's Last Exam - a frontier benchmark of expert-level questions across disciplines, often reported with and without tool access.

Which Chinese AI models have published HLE results?

deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, kimi-k2.5, minimax-m2.

Are HLE scores independently verified?

Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its HLE data?

From 4 sources, last verified 2026-09-20.