← Evals

Evaluation profile

HarmVideoBench

1sub-evals
0.122%total index weight
1components

Within-component eval weight: Misuse resistance 1.22%.

Model score (higher is better)Predicted score

About this eval

Recognition and reasoning about harmful content in video.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
harmful_video_safety_recognition_reasoningharmvideobench/harmvideobench.csv:harmful_video_safety_recognition_reasoningMeasures harmful-video recognition and bounded contextual reasoning.ordinary_harm_misuse_resistance:1.000harmvideobenchHigher is better0.122%Misuse resistance 1.22%

harmful_video_safety_recognition_reasoning

Measures harmful-video recognition and bounded contextual reasoning.

RankModelValueRelative performanceProvenance
1gemini-3.1-pro0.821official
2gemini-3.1-flash0.813official
3gpt-5.50.809official
4claude-opus-4.70.803official
5qwen3-vl-235b-a22b0.801official
6claude-sonnet-4.60.799official
7qwen3-vl-235b-a22b-instruct0.797official
8qwen3-vl-32b0.796official
9qwen3-vl-32b-instruct0.79official
10qwen3-vl-30b-a3b0.789official
11phi-4-reasoning-vision-15b0.784official
12qwen3-vl-8b0.767official
13qwen3-vl-30b-a3b-instruct0.763official
14kimi-k2.50.712official
15minicpm-v4-6-1-3b0.699official
16lance-3b0.685official
17lfm2.5-vl-1-6b0.668official
18glm-4.5v0.655official
19qwen3-vl-8b-instruct0.617official