← Evals

Evaluation profile

UAVBench safety-critical decision recognition

1sub-evals
0.512%total index weight
1components

Within-component eval weight: Responsible agency 3.41%.

Model score (higher is better)Predicted score

About this eval

Recognition of ethical and safety-critical UAV decisions in multiple-choice scenarios.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
ethical_safety_critical_accuracyuavbench/uavbench-style8.csv:ethical_safety_critical_accuracyMeasures whether the model chooses correct actions in unmanned-aircraft scenarios involving ethical and safety-critical decisions.responsible_agency_control:1.000uavbench-mcq-v1-2kHigher is better0.512%Responsible agency 3.41%

ethical_safety_critical_accuracy

Measures whether the model chooses correct actions in unmanned-aircraft scenarios involving ethical and safety-critical decisions.

RankModelValueRelative performanceProvenance
1deepseek-v3.2-exp0.775official
2gpt-50.76official
2qwen3-max0.76official
4deepseek-v30.755official
4qwen3-235b-a22b0.755official
6mistral-medium-3.10.75official
7phi-4-reasoning-plus0.735official
8gpt-4o0.73official
9deepseek-v3.1-terminus0.72official
9gemma-3n-e4b-it0.72official
9internvl3-78b0.72official
12gemini-2.5-flash0.715official
13gpt-4.10.7official
14grok-4-fast0.695official
15kimi-k20.685official
16ernie-4.5-300b-a47b0.68official
17claude-haiku-4.50.67official
17gpt-4.1-mini0.67official
19llama-4-scout0.635official
20qwen3-vl-8b-instruct0.62official
21olmo-2-0325-32b-instruct0.615official
22lfm2-8b-a1b0.575official
22lfm2.2-6b0.575official
24llama-3.1-8b-instruct0.57official
25granite-4.0-micro0.54official
26qwen-2.5-7b-instruct0.535official
27llama-3.2-3b-instruct0.475official