Evaluation profile
XSTest
2sub-evals
0.179%Safety weight
0.463%Freedom weight
2components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Two-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| safe_full_compliance_ratexstest/xstest.csv:safe_full_compliance_rateMeasures whether the model fully answers safe prompts that contain words often associated with harmful content. | Safety: benign_helpfulness_non_overrefusal:1.000 · Freedom: user_direction_following:1.000xstest | Safety: higher · Freedom: higher | 0.157% | 0.232% |
| unsafe_full_refusal_ratexstest/xstest.csv:unsafe_full_refusal_rateMeasures whether the model fully refuses matched prompts that actually request harmful assistance. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000xstest | Safety: higher · Freedom: lower | 0.0215% | 0.232% |
safe_full_compliance_rate
Measures whether the model fully answers safe prompts that contain words often associated with harmful content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | mistral-7b-instruct | 0.984 | official | |
| 2 | gpt-4 | 0.916 | official | |
| 3 | llama-2-70b-chat | 0.704 | official |
unsafe_full_refusal_rate
Measures whether the model fully refuses matched prompts that actually request harmful assistance.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4 | 0.975 | official | |
| 1 | llama-2-70b-chat | 0.975 | official | |
| 3 | mistral-7b-instruct | 0.235 | official |