Benchmarks / CyberGym
CyberGym
Cybersecurity agent benchmark focused on vulnerability discovery tasks.
Cybersecurity agent benchmark focused on vulnerability discovery tasks. The table below lists 3 recorded evaluations across 3 models: deepseek-v4-1-flash, deepseek-v4-pro, glm-5.3. All entries are labeled by source type (vendor_reported) with links to the original publication. See the Limitations section for comparability caveats before citing any score.
Benchmark methodology
| Task type | Cybersecurity vulnerability analysis (real-world vulnerability discovery) |
|---|---|
| Dataset size | Large-scale task suite sourced from ARVO and OSS-Fuzz (~240GB data) |
| Evaluation method | Docker-isolated environments; agents analyze vulnerabilities and generate proofs of concept; pre-/post-patch versions |
| Scoring | Success rate on vulnerability analysis tasks (PoC generation) |
Results
| Model | Model version | Score | Metric | Date | Source type | Source |
|---|---|---|---|---|---|---|
| deepseek-v4-1-flash | — | 88.1 | accuracy | 2026-09-10 | vendor_reported | link |
| deepseek-v4-pro | — | 83.3 | accuracy | 2026-08-13 | vendor_reported | link |
| glm-5.3 | vuln discovery (GLM-5.2: 77.2) | 84.5 | accuracy | 2026-08-18 | vendor_reported | link |
Limitations
All scores are vendor-reported and not independently verified.
Vendor-reported scores are labeled as such. Scores from incompatible benchmark versions are never mixed without explanation.
Sources
What does CyberGym measure?
Cybersecurity agent benchmark focused on vulnerability discovery tasks.
Which Chinese AI models have published CyberGym results?
deepseek-v4-1-flash, deepseek-v4-pro, glm-5.3.
Are CyberGym scores independently verified?
Results on this page are labeled by source type (vendor_reported). Vendor-reported scores are labeled as such, and scores from incompatible benchmark versions are never mixed.
Where does China AI Hub get its CyberGym data?
From 3 sources, last verified 2026-09-20.