Benchmarks / MMMU-Pro
MMMU-Pro
Multimodal, multi-discipline understanding benchmark with college-level questions requiring reasoning.
Multimodal, multi-discipline understanding benchmark with college-level questions requiring reasoning. The table below lists 2 recorded evaluations across 2 models: kimi-k3, kimi-k2.5. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.
Benchmark methodology
| Task type | Multimodal understanding and reasoning (college-level, multi-discipline) |
|---|---|
| Dataset size | 1,730 questions in standard format plus 1,730 vision-augmented variants (3,460 total); parent MMMU = 11.5K questions across 6 disciplines, 30 subjects |
| Evaluation method | Multiple-choice questions with interleaved images; vision-only input setting removes text leakage |
| Scoring | Accuracy (% correct answers) |
Results
| Model | Model version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| kimi-k3 | — | 81.6 (83.4 with tools) | accuracy | 2026-07 | vendor_reported | link |
| kimi-k2.5 | — | 78.5 | accuracy | — | vendor_reported | link |
Limitations
All scores are vendor-reported and not independently verified. With-tools and without-tools results are not directly comparable.
Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.
Sources
What does MMMU-Pro measure?
Multimodal, multi-discipline understanding benchmark with college-level questions requiring reasoning.
Which Chinese AI models have published MMMU-Pro results?
kimi-k3, kimi-k2.5.
Are MMMU-Pro scores independently verified?
Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.
Where does China AI Hub get its MMMU-Pro data?
From 3 sources, last verified 2026-09-20.