Evaluation profile
Qwen2 Safety Panel
4sub-evals
0.146%Safety weight
0.463%Freedom weight
2components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Harmful-request and privacy refusal defects
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| fraudqwen2-system-card/qwen2-system-card.csv:fraudMeasures harmful-response rate for fraud requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000qwen2-safety-panel | Safety: lower · Freedom: higher | 0.017% | 0.116% |
| illegalqwen2-system-card/qwen2-system-card.csv:illegalMeasures harmful-response rate for illegal requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000qwen2-safety-panel | Safety: lower · Freedom: higher | 0.017% | 0.116% |
| pornographyqwen2-system-card/qwen2-system-card.csv:pornographyMeasures harmful-response rate for pornographic requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000qwen2-safety-panel | Safety: lower · Freedom: higher | 0.0136% | 0.116% |
| privacyqwen2-system-card/qwen2-system-card.csv:privacyMeasures harmful-response rate for privacy-violating requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000qwen2-safety-panel | Safety: lower · Freedom: higher | 0.0987% | 0.116% |
fraud
Measures harmful-response rate for fraud requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen2-72b-instruct | 2.41 | official | |
| 2 | gpt-4 | 6.87 | official | |
| 3 | mixtral-8x22b-instruct | 8.49 | official |
illegal
Measures harmful-response rate for illegal requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4 | 0 | official | |
| 1 | qwen2-72b-instruct | 0 | official | |
| 3 | mixtral-8x22b-instruct | 6.87 | official |
pornography
Measures harmful-response rate for pornographic requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen2-72b-instruct | 22.91 | official | |
| 2 | gpt-4 | 23.63 | official | |
| 3 | mixtral-8x22b-instruct | 33.82 | official |
privacy
Measures harmful-response rate for privacy-violating requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen2-72b-instruct | 2.47 | official | |
| 2 | gpt-4 | 3.37 | official | |
| 3 | mixtral-8x22b-instruct | 15.03 | official |