China AI Hub China AI Hub

Benchmarks / Video-MME

Video-MME

Video understanding benchmark spanning various video durations and domains.

Video-MME
Image: AI-generated illustration (Seedream)

Video understanding benchmark spanning various video durations and domains. The table below lists 2 recorded evaluations across 2 models: kimi-k3, kimi-k2.5. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.

Benchmark methodology

Methodology of Video-MME
Task type Video understanding (multimodal video analysis)
Dataset size 900 videos (254 hours total), 2,700 human-annotated question-answer pairs
Evaluation method Video QA with subtitles and audio modalities; duration-stratified evaluation
Scoring Accuracy (% correct); variants with/without subtitles
Last verified: · Data status: Current · Next review:

Results

Results for Video-MME
Model Model version Score Metric Date Source type Source
kimi-k3 with subtitles 90 accuracy 2026-07 vendor_reported link
kimi-k2.5 87.4 accuracy vendor_reported link

Limitations

All scores are vendor-reported and not independently verified. Subtitle usage differs between evaluations.

Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.

Sources

What does Video-MME measure?

Video understanding benchmark spanning various video durations and domains.

Which Chinese AI models have published Video-MME results?

kimi-k3, kimi-k2.5.

Are Video-MME scores independently verified?

Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.

Where does China AI Hub get its Video-MME data?

From 3 sources, last verified 2026-09-20.