← Evals

Evaluation profile

Qwen2 Safety Panel

4sub-evals
0.146%Safety weight
0.463%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Harmful-request and privacy refusal defects

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
fraudqwen2-system-card/qwen2-system-card.csv:fraudMeasures harmful-response rate for fraud requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000qwen2-safety-panelSafety: lower · Freedom: higher0.017%0.116%
illegalqwen2-system-card/qwen2-system-card.csv:illegalMeasures harmful-response rate for illegal requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000qwen2-safety-panelSafety: lower · Freedom: higher0.017%0.116%
pornographyqwen2-system-card/qwen2-system-card.csv:pornographyMeasures harmful-response rate for pornographic requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000qwen2-safety-panelSafety: lower · Freedom: higher0.0136%0.116%
privacyqwen2-system-card/qwen2-system-card.csv:privacyMeasures harmful-response rate for privacy-violating requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000qwen2-safety-panelSafety: lower · Freedom: higher0.0987%0.116%

fraud

Measures harmful-response rate for fraud requests.

RankModelValueRelative performanceProvenance
1qwen2-72b-instruct2.41official
2gpt-46.87official
3mixtral-8x22b-instruct8.49official

illegal

Measures harmful-response rate for illegal requests.

RankModelValueRelative performanceProvenance
1gpt-40official
1qwen2-72b-instruct0official
3mixtral-8x22b-instruct6.87official

pornography

Measures harmful-response rate for pornographic requests.

RankModelValueRelative performanceProvenance
1qwen2-72b-instruct22.91official
2gpt-423.63official
3mixtral-8x22b-instruct33.82official

privacy

Measures harmful-response rate for privacy-violating requests.

RankModelValueRelative performanceProvenance
1qwen2-72b-instruct2.47official
2gpt-43.37official
3mixtral-8x22b-instruct15.03official