Benchmarks / Video-MME
Video-MME
Video understanding benchmark spanning various video durations and domains.
Video understanding benchmark spanning various video durations and domains. The table below lists 2 recorded evaluations across 2 models: kimi-k3, kimi-k2.5. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.
Benchmark methodology
| Task type | Video understanding (multimodal video analysis) |
|---|---|
| Dataset size | 900 videos (254 hours total), 2,700 human-annotated question-answer pairs |
| Evaluation method | Video QA with subtitles and audio modalities; duration-stratified evaluation |
| Scoring | Accuracy (% correct); variants with/without subtitles |
Results
| Model | Model version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| kimi-k3 | with subtitles | 90 | accuracy | 2026-07 | vendor_reported | link |
| kimi-k2.5 | — | 87.4 | accuracy | — | vendor_reported | link |
Limitations
All scores are vendor-reported and not independently verified. Subtitle usage differs between evaluations.
Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.
Sources
What does Video-MME measure?
Video understanding benchmark spanning various video durations and domains.
Which Chinese AI models have published Video-MME results?
kimi-k3, kimi-k2.5.
Are Video-MME scores independently verified?
Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.
Where does China AI Hub get its Video-MME data?
From 3 sources, last verified 2026-09-20.