Evaluation profile
FLAMES
Within-component eval weight: Human rights 1.89% · Fairness 2.29% · Misuse resistance 1.38%.
About this eval
Chinese fairness, safety, morality, legality, and data-protection behavior on the FLAMES leaderboard.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| data_protectionflames/flames.csv:data_protectionMeasures whether the model avoids responses involving privacy violations and misuse of personal data in Chinese-language prompts. | human_rights_systemic_harm:1.000flames/flames.csv | Higher is better | 0.283% | Human rights 1.89% |
| fairnessflames/flames.csv:fairnessMeasures whether the model avoids responses involving discriminatory or unfair treatment in Chinese-language prompts. | fairness_nondiscrimination:1.000flames/flames.csv | Higher is better | 0.229% | Fairness 2.29% |
| legalityflames/flames.csv:legalityMeasures whether the model avoids responses involving illegal conduct in Chinese-language prompts. | ordinary_harm_misuse_resistance:1.000flames/flames.csv | Higher is better | 0.0461% | Misuse resistance 0.461% |
| moralityflames/flames.csv:moralityMeasures whether the model avoids responses involving socially or morally harmful conduct in Chinese-language prompts. | ordinary_harm_misuse_resistance:1.000flames/flames.csv | Higher is better | 0.0461% | Misuse resistance 0.461% |
| safetyflames/flames.csv:safetyMeasures whether the model avoids responses involving harmful instructions and unsafe conduct in Chinese-language prompts. | ordinary_harm_misuse_resistance:1.000flames/flames.csv | Higher is better | 0.0461% | Misuse resistance 0.461% |
data_protection
Measures whether the model avoids responses involving privacy violations and misuse of personal data in Chinese-language prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | internlm-chat-20b | 63.16 | official | |
| 2 | internlm-chat-7b | 61.84 | official | |
| 3 | qwen-14b-chat | 55.26 | official | |
| 4 | gpt-4 | 50 | official | |
| 5 | ernie-bot | 46.05 | official | |
| 6 | chatglm2-6b | 43.42 | official | |
| 6 | qwen-7b-chat | 43.42 | official | |
| 8 | baichuan-2-7b-chat | 40.79 | official | |
| 9 | baichuan-2-13b-chat | 39.47 | official | |
| 10 | chatglm3-6b | 38.16 | official | |
| 11 | chatglm-6b | 32.89 | official | |
| 11 | moss-16b | 32.89 | official | |
| 13 | belle-13b | 26.32 | official |
fairness
Measures whether the model avoids responses involving discriminatory or unfair treatment in Chinese-language prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | internlm-chat-20b | 52.61 | official | |
| 2 | internlm-chat-7b | 44.58 | official | |
| 3 | ernie-bot | 42.97 | official | |
| 4 | baichuan-2-7b-chat | 42.17 | official | |
| 5 | gpt-4 | 41.37 | official | |
| 6 | baichuan-2-13b-chat | 38.55 | official | |
| 7 | chatglm3-6b | 37.75 | official | |
| 8 | qwen-7b-chat | 36.14 | official | |
| 9 | moss-16b | 33.33 | official | |
| 10 | chatglm2-6b | 31.73 | official | |
| 11 | qwen-14b-chat | 30.92 | official | |
| 12 | chatglm-6b | 26.91 | official | |
| 13 | belle-13b | 22.09 | official |
legality
Measures whether the model avoids responses involving illegal conduct in Chinese-language prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | internlm-chat-7b | 76.09 | official | |
| 2 | internlm-chat-20b | 71.74 | official | |
| 3 | ernie-bot | 60.87 | official | |
| 4 | baichuan-2-7b-chat | 52.17 | official | |
| 5 | chatglm-6b | 50 | official | |
| 5 | moss-16b | 50 | official | |
| 7 | baichuan-2-13b-chat | 39.13 | official | |
| 7 | belle-13b | 39.13 | official | |
| 9 | qwen-14b-chat | 32.61 | official | |
| 10 | gpt-4 | 30.43 | official | |
| 10 | qwen-7b-chat | 30.43 | official | |
| 12 | chatglm2-6b | 28.26 | official | |
| 12 | chatglm3-6b | 28.26 | official |
morality
Measures whether the model avoids responses involving socially or morally harmful conduct in Chinese-language prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | internlm-chat-20b | 54.23 | official | |
| 1 | qwen-14b-chat | 54.23 | official | |
| 3 | internlm-chat-7b | 51.24 | official | |
| 4 | gpt-4 | 50.75 | official | |
| 5 | ernie-bot | 47.76 | official | |
| 6 | baichuan-2-13b-chat | 44.78 | official | |
| 6 | chatglm3-6b | 44.78 | official | |
| 8 | chatglm2-6b | 43.28 | official | |
| 9 | chatglm-6b | 40.3 | official | |
| 9 | qwen-7b-chat | 40.3 | official | |
| 11 | baichuan-2-7b-chat | 39.3 | official | |
| 12 | moss-16b | 31.34 | official | |
| 13 | belle-13b | 20.9 | official |
safety
Measures whether the model avoids responses involving harmful instructions and unsafe conduct in Chinese-language prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | baichuan-2-7b-chat | 56.41 | official | |
| 2 | baichuan-2-13b-chat | 53.85 | official | |
| 3 | internlm-chat-20b | 51.05 | official | |
| 4 | qwen-14b-chat | 36.83 | official | |
| 5 | internlm-chat-7b | 35.9 | official | |
| 6 | moss-16b | 33.33 | official | |
| 7 | chatglm3-6b | 32.63 | official | |
| 8 | ernie-bot | 32.17 | official | |
| 9 | qwen-7b-chat | 31.93 | official | |
| 10 | gpt-4 | 27.51 | official | |
| 11 | chatglm2-6b | 22.61 | official | |
| 12 | belle-13b | 15.38 | official | |
| 12 | chatglm-6b | 15.38 | official |