Evaluation profile
CValues
1sub-evals
0.152%total index weight
1components
Within-component eval weight: Misuse resistance 1.52%.
Model score (higher is better)Predicted score
About this eval
Chinese value alignment and social-norm understanding.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| scorecvalues/cvalues.csv:scoreMeasures whether the model chooses safe, value-aligned responses to Chinese prompts about socially sensitive conduct. | ordinary_harm_misuse_resistance:1.000cvalues/cvalues.csv | Higher is better | 0.152% | Misuse resistance 1.52% |
score
Measures whether the model chooses safe, value-aligned responses to Chinese prompts about socially sensitive conduct.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | chatgpt | 96.9 | official | |
| 2 | chinese-alpaca-plus-7b | 95.3 | official | |
| 3 | chatglm-6b | 95 | official | |
| 4 | chatplug-13b | 94.7 | official | |
| 5 | chinese-alpaca-plus-13b | 93 | official | |
| 6 | moss | 88.9 | official | |
| 7 | chatplug-3.7b | 88.8 | official | |
| 8 | ziya-llama-13b-v1 | 77.8 | official | |
| 9 | belle-7b-2m | 72.8 | official | |
| 10 | chinese-llama-13b | 53 | official |