← Evals

Evaluation profile

AgentDrive Safety Compliance

1sub-evals
0.194%total index weight
1components

Within-component eval weight: Misuse resistance 1.94%.

Model score (higher is better)Predicted score

About this eval

Policy and scenario safety knowledge for autonomous-system decisions.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scragentdrive/agentdrive.csv:scrMeasures textual policy/scenario safety knowledge.ordinary_harm_misuse_resistance:1.000agentdrive-safety-complianceHigher is better0.194%Misuse resistance 1.94%

scr

Measures textual policy/scenario safety knowledge.

RankModelValueRelative performanceProvenance
1chatgpt-4o97.5official
2ernie-4.5-300b-a47b96.25official
2gpt-4.1-mini96.25official
2gpt-596.25official
2mistral-medium-3.196.25official
6deepseek-v395official
6phi-4-reasoning-plus95official
6qwen3-max95official
6qwen3-next-80b-a3b-instruct95official
10claude-opus-4.193.75official
10gemini-2.5-flash93.75official
10gpt-4.1-nano93.75official
13gpt-4.192.5official
13grok-3-mini92.5official
13qwen3-235b-a22b92.5official
16claude-sonnet-4.590official
16gpt-4o90official
16kimi-dev-72b90official
19kimi-k289.38official
20deepseek-v3.1-terminus88.75official
20grok-4-fast88.75official
20hermes-3-llama-3.1-405b88.75official
20nova-micro88.75official
24cogito-v2-preview-llama-109b-moe82.5official
24deepseek-v3.2-exp82.5official
24phi-3-medium-128k-instruct82.5official
24qwen-turbo82.5official
28gemma-3-12b77.5official
29olmo-2-0325-32b-instruct76.25official
30llama-3.1-70b-instruct75official
30llama-4-maverick75official
32claude-3.7-sonnet68.75official
33gemma-3-27b-it66.25official
34afm-4.5b61.25official
34llama-3.3-70b-instruct61.25official
36gpt-4o-mini58.75official
36llama-4-scout58.75official
38command-r7b-12-202457.5official
38wizardlm-2-8x22b57.5official
40qwen-2.5-7b-instruct53.75official
41lfm-3b52.5official
42llama-3.1-8b-instruct50official
42mistral-nemo50official
44gpt-3.5-turbo-instruct43.75official
44llama-3.2-3b-instruct43.75official
46glm-4.641.25official
47llama-3.2-1b-instruct40official
48gpt-oss-120b10official