Benchmarks / HLE
HLE
Humanity's Last Exam - a frontier benchmark of expert-level questions across disciplines, often reported with and without tool access.
Humanity’s Last Exam - a frontier benchmark of expert-level questions across disciplines, often reported with and without tool access. The table below lists 6 recorded evaluations across 6 models: deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, kimi-k2.5, minimax-m2. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.
Benchmark methodology
| Task type | Expert-level academic Q&A across mathematics, humanities and natural sciences |
|---|---|
| Dataset size | 2,500 questions across dozens of subjects |
| Evaluation method | Multiple-choice and short-answer questions developed by subject-matter experts; multimodal; suitable for automated grading |
| Scoring | Accuracy (% correct); frequently reported with and without tool access |
| Contamination notes | Dataset includes a canary string (hle:3r2s:26b5c67b-...) to aid model builders in filtering the dataset from future training. |
Results
| Model | Model version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| deepseek-v4-1-flash | 39.1 on pure-text subset | 36.8 | accuracy | 2026-09-10 | vendor_reported | link |
| deepseek-v4-pro | — | 42.7 (60.0 with tools) | accuracy | 2026-08-13 | vendor_reported | link |
| qwen3.8-max | — | 43.6 (56.2 with tools) | accuracy | 2026-08 | vendor_reported | link |
| kimi-k3 | HLE-Full | 43.5 (56.0 with tools) | accuracy | 2026-07 | vendor_reported | link |
| kimi-k2.5 | HLE-Full | 30.1 (50.2 with tools) | accuracy | — | vendor_reported | link |
| minimax-m2 | — | 12.5 without tools / 31.8 with tools | accuracy | 2025-10 | vendor_reported | link |
Limitations
All scores are vendor-reported and not independently verified. With-tools and without-tools results are not directly comparable; the setting is recorded per score.
Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.
Sources
What does HLE measure?
Humanity's Last Exam - a frontier benchmark of expert-level questions across disciplines, often reported with and without tool access.
Which Chinese AI models have published HLE results?
deepseek-v4-1-flash, deepseek-v4-pro, qwen3.8-max, kimi-k3, kimi-k2.5, minimax-m2.
Are HLE scores independently verified?
Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.
Where does China AI Hub get its HLE data?
From 4 sources, last verified 2026-09-20.