Evaluation profile
OR-Bench
2sub-evals
1.03%Safety weight
2.68%Freedom weight
2components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Appropriate refusal calibration across seemingly toxic but safe requests and genuinely toxic requests.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| over_refusal_rateor-bench/or-bench.csv:over_refusal_rateMeasures whether the model incorrectly refuses seemingly sensitive but safe requests. | Safety: benign_helpfulness_non_overrefusal:1.000 · Freedom: user_direction_following:1.000or-bench/or-bench.csv | Safety: lower · Freedom: lower | 0.907% | 1.34% |
| toxic_acceptance_rateor-bench/or-bench.csv:toxic_acceptance_rateMeasures whether the model answers genuinely harmful requests instead of refusing them. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000or-bench/or-bench.csv | Safety: lower · Freedom: higher | 0.124% | 1.34% |
over_refusal_rate
Measures whether the model incorrectly refuses seemingly sensitive but safe requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o | 6.8 | official | |
| 2 | gemini-1.0-pro | 9.7 | official | |
| 3 | mistral-large | 9.8 | official | |
| 4 | gpt-4-turbo | 12.5 | official | |
| 5 | mistral-small | 13.3 | official | |
| 6 | mistral-medium | 14 | official | |
| 7 | gemma-7b-it | 26.3 | official | |
| 8 | gpt-3.5-turbo | 36.17 | official | |
| 9 | llama-3-70b-instruct | 37.7 | official | |
| 10 | qwen1.5-7b-chat | 39.2 | official | |
| 11 | claude-3.5-sonnet | 43.8 | official | |
| 12 | qwen1.5-72b-chat | 46.9 | official | |
| 13 | qwen1.5-32b-chat | 50.8 | official | |
| 14 | gemma-2-27b-it | 62 | official | |
| 15 | llama-3-8b-instruct | 69.4 | official | |
| 16 | gemma-2-9b-it | 79.9 | official | |
| 17 | gemini-1.5-flash | 84.3 | official | |
| 18 | llama-2-7b-chat | 87.5 | official | |
| 19 | gemini-1.5-pro | 88 | official | |
| 20 | claude-3-opus | 91 | official | |
| 20 | llama-2-13b-chat | 91 | official | |
| 22 | claude-3-sonnet | 94.5 | official | |
| 23 | llama-2-70b-chat | 96.1 | official | |
| 24 | claude-3-haiku | 96.3 | official | |
| 25 | claude-2.1 | 99.8 | official |
toxic_acceptance_rate
Measures whether the model answers genuinely harmful requests instead of refusing them.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2.1 | 0 | official | |
| 2 | claude-3-haiku | 0.3 | official | |
| 2 | claude-3-sonnet | 0.3 | official | |
| 2 | llama-2-13b-chat | 0.3 | official | |
| 2 | llama-2-70b-chat | 0.3 | official | |
| 6 | llama-2-7b-chat | 0.4 | official | |
| 7 | gemini-1.5-pro | 0.6 | official | |
| 8 | gemini-1.5-flash | 1.2 | official | |
| 9 | gemma-2-9b-it | 1.5 | official | |
| 10 | claude-3-opus | 1.9 | official | |
| 11 | gemma-2-27b-it | 3.2 | official | |
| 12 | claude-3.5-sonnet | 3.4 | official | |
| 13 | qwen1.5-32b-chat | 4.4 | official | |
| 14 | llama-3-8b-instruct | 5 | official | |
| 15 | gpt-4-turbo | 5.25 | official | |
| 16 | qwen1.5-72b-chat | 5.6 | official | |
| 17 | gemma-7b-it | 14.5 | official | |
| 18 | qwen1.5-7b-chat | 15 | official | |
| 19 | gpt-4o | 15.1 | official | |
| 20 | gpt-3.5-turbo | 17.03 | official | |
| 21 | mistral-small | 20.3 | official | |
| 22 | gemini-1.0-pro | 21.3 | official | |
| 22 | llama-3-70b-instruct | 21.3 | official | |
| 24 | mistral-medium | 22.5 | official | |
| 25 | mistral-large | 27.2 | official |