← Evals

Evaluation profile

SafetyPrompts

1sub-evals
0.233%Safety weight
0.691%Freedom weight
4components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Safety performance on Chinese safety prompts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
scoresafety-prompts/safety-prompts.csv:scoreMeasures whether responses avoid offensive, biased, illegal, physically harmful, psychologically harmful, privacy-invasive, and unethical content under varied instruction attacks.Safety: human_rights_systemic_harm:0.214|fairness_nondiscrimination:0.071|truthfulness_honesty:0.071|ordinary_harm_misuse_resistance:0.643 · Freedom: user_direction_following:1.000safety-prompts/safety-prompts.csvSafety: higher · Freedom: lower0.233%0.691%

score

Measures whether responses avoid offensive, biased, illegal, physically harmful, psychologically harmful, privacy-invasive, and unethical content under varied instruction attacks.

RankModelValueRelative performanceProvenance
1gpt-3.5-turbo98.37official
2chatglm-130b96.91official
3chatglm-6b96.81official
4minichat-100b95.25official
5minichat-10b92.18official
6belle86.14official
7text-davinci-00384.76official
8text-davinci-00259.89official
9text-davinci-00157.92official
10chatyuan48.9official
11opd46.74official
12davinci42.98official
13bloom41.65official
14eva2.021.68official
15eva1.019.7official