← Evals

Evaluation profile

Agent-SafetyBench

8sub-evals
0.911%total index weight
2components

Within-component eval weight: Responsible agency 5.83% · Misuse resistance 0.355%.

Model score (higher is better)Predicted score

About this eval

Safety of LLM agents in risky tool-use or action settings.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
compromise_availabilityagent-safetybench/agent-safetybench.csv:compromise_availabilityMeasures whether a tool-using agent avoids actions that disrupt system or service availability.responsible_agency_control:1.000agent-safetybench/agent-safetybench.csvHigher is better0.131%Responsible agency 0.875%
harmful_vulnerable_codeagent-safetybench/agent-safetybench.csv:harmful_vulnerable_codeMeasures whether a tool-using agent avoids writing harmful or vulnerable code.responsible_agency_control:1.000agent-safetybench/agent-safetybench.csvHigher is better0.131%Responsible agency 0.875%
leak_sensitive_informationagent-safetybench/agent-safetybench.csv:leak_sensitive_informationMeasures whether a tool-using agent avoids leaking sensitive information.responsible_agency_control:1.000agent-safetybench/agent-safetybench.csvHigher is better0.131%Responsible agency 0.875%
physical_harmagent-safetybench/agent-safetybench.csv:physical_harmMeasures whether a tool-using agent avoids actions that could cause physical harm.responsible_agency_control:1.000agent-safetybench/agent-safetybench.csvHigher is better0.131%Responsible agency 0.875%
produce_unsafe_informationagent-safetybench/agent-safetybench.csv:produce_unsafe_informationMeasures whether a tool-using agent avoids directly producing unsafe information.ordinary_harm_misuse_resistance:1.000agent-safetybench/agent-safetybench.csvHigher is better0.0355%Misuse resistance 0.355%
property_lossagent-safetybench/agent-safetybench.csv:property_lossMeasures whether a tool-using agent avoids actions that could cause property loss.responsible_agency_control:1.000agent-safetybench/agent-safetybench.csvHigher is better0.131%Responsible agency 0.875%
spread_unsafe_informationagent-safetybench/agent-safetybench.csv:spread_unsafe_informationMeasures whether a tool-using agent avoids spreading unsafe information through tools.responsible_agency_control:1.000agent-safetybench/agent-safetybench.csvHigher is better0.131%Responsible agency 0.875%
violate_law_ethicsagent-safetybench/agent-safetybench.csv:violate_law_ethicsMeasures whether a tool-using agent avoids actions that violate legal or ethical constraints.responsible_agency_control:1.000agent-safetybench/agent-safetybench.csvHigher is better0.0875%Responsible agency 0.583%

compromise_availability

Measures whether a tool-using agent avoids actions that disrupt system or service availability.

RankModelValueRelative performanceProvenance
1glm-4-9b-chat50.8official
2claude-3-opus43.2official
3claude-3.5-sonnet40.4official
4gpt-4-turbo37.6official
5gpt-4o35.2official
5qwen2.5-72b-instruct35.2official
7deepseek-v2.533.2official
8gemini-1.5-pro30.8official
9gemini-1.5-flash30official
10qwen-2.5-14b-instruct29.2official
11claude-3.5-haiku26.4official
12llama-3.1-70b-instruct24official
13gpt-4o-mini23.6official
14llama-3.1-405b-instruct19.6official
15qwen-2.5-7b-instruct17.2official
16llama-3.1-8b-instruct12.8official

harmful_vulnerable_code

Measures whether a tool-using agent avoids writing harmful or vulnerable code.

RankModelValueRelative performanceProvenance
1claude-3.5-sonnet64.8official
2claude-3.5-haiku60.8official
3claude-3-opus60official
4gemini-1.5-flash48.4official
5gemini-1.5-pro42official
6llama-3.1-405b-instruct40.4official
7gpt-4-turbo38.4official
8gpt-4o35.6official
9deepseek-v2.530.4official
10llama-3.1-70b-instruct29.6official
10qwen2.5-72b-instruct29.6official
12qwen-2.5-14b-instruct29.2official
13gpt-4o-mini25.2official
14llama-3.1-8b-instruct24.8official
15glm-4-9b-chat23.2official
16qwen-2.5-7b-instruct10.8official

leak_sensitive_information

Measures whether a tool-using agent avoids leaking sensitive information.

RankModelValueRelative performanceProvenance
1claude-3-opus60.4official
2claude-3.5-sonnet57.6official
3claude-3.5-haiku47.2official
4gpt-4o44.4official
5gemini-1.5-flash39.2official
6glm-4-9b-chat38.4official
7gpt-4-turbo36.8official
8qwen2.5-72b-instruct32.8official
9deepseek-v2.531.2official
10gemini-1.5-pro30official
11gpt-4o-mini28official
12llama-3.1-405b-instruct25.2official
13qwen-2.5-14b-instruct24.4official
14llama-3.1-70b-instruct20official
15qwen-2.5-7b-instruct13.2official
16llama-3.1-8b-instruct10official

physical_harm

Measures whether a tool-using agent avoids actions that could cause physical harm.

RankModelValueRelative performanceProvenance
1claude-3.5-sonnet69.6official
2claude-3-opus61.6official
3gpt-4o53.2official
4claude-3.5-haiku45.6official
5glm-4-9b-chat41.6official
6gemini-1.5-flash38.8official
6gpt-4-turbo38.8official
8deepseek-v2.534.4official
9qwen2.5-72b-instruct29.6official
10gemini-1.5-pro28.8official
11qwen-2.5-14b-instruct28official
12gpt-4o-mini24.4official
12llama-3.1-405b-instruct24.4official
14llama-3.1-70b-instruct23.2official
15qwen-2.5-7b-instruct17.6official
16llama-3.1-8b-instruct11.2official

produce_unsafe_information

Measures whether a tool-using agent avoids directly producing unsafe information.

RankModelValueRelative performanceProvenance
1claude-3-opus100official
1claude-3.5-haiku100official
1claude-3.5-sonnet100official
4llama-3.1-405b-instruct98.8official
5gpt-4o-mini98.4official
6qwen2.5-72b-instruct97.2official
7gpt-4o95.6official
8gpt-4-turbo94.4official
9llama-3.1-70b-instruct93.2official
10gemini-1.5-pro84.8official
11gemini-1.5-flash82.4official
12qwen-2.5-14b-instruct81.2official
13deepseek-v2.576.8official
14llama-3.1-8b-instruct74.8official
15qwen-2.5-7b-instruct57.6official
16glm-4-9b-chat57.2official

property_loss

Measures whether a tool-using agent avoids actions that could cause property loss.

RankModelValueRelative performanceProvenance
1claude-3-opus60.4official
2claude-3.5-sonnet58.4official
3gpt-4o48.4official
4glm-4-9b-chat48official
5claude-3.5-haiku46official
6gpt-4-turbo43.2official
7gemini-1.5-flash41.6official
8qwen2.5-72b-instruct38official
9gemini-1.5-pro37.6official
10deepseek-v2.536.8official
11qwen-2.5-14b-instruct31.2official
12gpt-4o-mini30official
13llama-3.1-70b-instruct28.4official
14llama-3.1-405b-instruct27.6official
15qwen-2.5-7b-instruct15.6official
16llama-3.1-8b-instruct12.4official

spread_unsafe_information

Measures whether a tool-using agent avoids spreading unsafe information through tools.

RankModelValueRelative performanceProvenance
1claude-3-opus35.6official
2claude-3.5-haiku33.6official
3claude-3.5-sonnet32.4official
4gemini-1.5-flash20.8official
5gemini-1.5-pro18.8official
6llama-3.1-405b-instruct14.4official
7gpt-4-turbo12.4official
7gpt-4o12.4official
9qwen2.5-72b-instruct12official
10qwen-2.5-14b-instruct11.2official
11llama-3.1-70b-instruct10.8official
12deepseek-v2.58.8official
13qwen-2.5-7b-instruct7.6official
14gpt-4o-mini6.8official
15llama-3.1-8b-instruct6.4official
16glm-4-9b-chat6official

violate_law_ethics

Measures whether a tool-using agent avoids actions that violate legal or ethical constraints.

RankModelValueRelative performanceProvenance
1claude-3-opus56.8official
2claude-3.5-sonnet52official
3claude-3.5-haiku41.2official
4gpt-4-turbo33.2official
5llama-3.1-405b-instruct32.8official
6gemini-1.5-flash32official
7gpt-4o28.8official
8glm-4-9b-chat27.2official
9gemini-1.5-pro26.8official
10qwen2.5-72b-instruct24official
11deepseek-v2.522official
12llama-3.1-70b-instruct20.4official
12qwen-2.5-14b-instruct20.4official
14gpt-4o-mini13.2official
15qwen-2.5-7b-instruct10.4official
16llama-3.1-8b-instruct6.8official