← Evals

Evaluation profile

OR-Bench

2sub-evals
1.03%Safety weight
2.68%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Appropriate refusal calibration across seemingly toxic but safe requests and genuinely toxic requests.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
over_refusal_rateor-bench/or-bench.csv:over_refusal_rateMeasures whether the model incorrectly refuses seemingly sensitive but safe requests.Safety: benign_helpfulness_non_overrefusal:1.000 · Freedom: user_direction_following:1.000or-bench/or-bench.csvSafety: lower · Freedom: lower0.907%1.34%
toxic_acceptance_rateor-bench/or-bench.csv:toxic_acceptance_rateMeasures whether the model answers genuinely harmful requests instead of refusing them.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000or-bench/or-bench.csvSafety: lower · Freedom: higher0.124%1.34%

over_refusal_rate

Measures whether the model incorrectly refuses seemingly sensitive but safe requests.

RankModelValueRelative performanceProvenance
1gpt-4o6.8official
2gemini-1.0-pro9.7official
3mistral-large9.8official
4gpt-4-turbo12.5official
5mistral-small13.3official
6mistral-medium14official
7gemma-7b-it26.3official
8gpt-3.5-turbo36.17official
9llama-3-70b-instruct37.7official
10qwen1.5-7b-chat39.2official
11claude-3.5-sonnet43.8official
12qwen1.5-72b-chat46.9official
13qwen1.5-32b-chat50.8official
14gemma-2-27b-it62official
15llama-3-8b-instruct69.4official
16gemma-2-9b-it79.9official
17gemini-1.5-flash84.3official
18llama-2-7b-chat87.5official
19gemini-1.5-pro88official
20claude-3-opus91official
20llama-2-13b-chat91official
22claude-3-sonnet94.5official
23llama-2-70b-chat96.1official
24claude-3-haiku96.3official
25claude-2.199.8official

toxic_acceptance_rate

Measures whether the model answers genuinely harmful requests instead of refusing them.

RankModelValueRelative performanceProvenance
1claude-2.10official
2claude-3-haiku0.3official
2claude-3-sonnet0.3official
2llama-2-13b-chat0.3official
2llama-2-70b-chat0.3official
6llama-2-7b-chat0.4official
7gemini-1.5-pro0.6official
8gemini-1.5-flash1.2official
9gemma-2-9b-it1.5official
10claude-3-opus1.9official
11gemma-2-27b-it3.2official
12claude-3.5-sonnet3.4official
13qwen1.5-32b-chat4.4official
14llama-3-8b-instruct5official
15gpt-4-turbo5.25official
16qwen1.5-72b-chat5.6official
17gemma-7b-it14.5official
18qwen1.5-7b-chat15official
19gpt-4o15.1official
20gpt-3.5-turbo17.03official
21mistral-small20.3official
22gemini-1.0-pro21.3official
22llama-3-70b-instruct21.3official
24mistral-medium22.5official
25mistral-large27.2official