Evaluation profile
HarmVideoBench
1sub-evals
0.122%total index weight
1components
Within-component eval weight: Misuse resistance 1.22%.
Model score (higher is better)Predicted score
About this eval
Recognition and reasoning about harmful content in video.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| harmful_video_safety_recognition_reasoningharmvideobench/harmvideobench.csv:harmful_video_safety_recognition_reasoningMeasures harmful-video recognition and bounded contextual reasoning. | ordinary_harm_misuse_resistance:1.000harmvideobench | Higher is better | 0.122% | Misuse resistance 1.22% |
harmful_video_safety_recognition_reasoning
Measures harmful-video recognition and bounded contextual reasoning.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-3.1-pro | 0.821 | official | |
| 2 | gemini-3.1-flash | 0.813 | official | |
| 3 | gpt-5.5 | 0.809 | official | |
| 4 | claude-opus-4.7 | 0.803 | official | |
| 5 | qwen3-vl-235b-a22b | 0.801 | official | |
| 6 | claude-sonnet-4.6 | 0.799 | official | |
| 7 | qwen3-vl-235b-a22b-instruct | 0.797 | official | |
| 8 | qwen3-vl-32b | 0.796 | official | |
| 9 | qwen3-vl-32b-instruct | 0.79 | official | |
| 10 | qwen3-vl-30b-a3b | 0.789 | official | |
| 11 | phi-4-reasoning-vision-15b | 0.784 | official | |
| 12 | qwen3-vl-8b | 0.767 | official | |
| 13 | qwen3-vl-30b-a3b-instruct | 0.763 | official | |
| 14 | kimi-k2.5 | 0.712 | official | |
| 15 | minicpm-v4-6-1-3b | 0.699 | official | |
| 16 | lance-3b | 0.685 | official | |
| 17 | lfm2.5-vl-1-6b | 0.668 | official | |
| 18 | glm-4.5v | 0.655 | official | |
| 19 | qwen3-vl-8b-instruct | 0.617 | official |