Evaluation profile
S-Eval
1sub-evals
0.225%total index weight
1components
Within-component eval weight: Misuse resistance 2.25%.
Model score (higher is better)Predicted score
About this eval
Automated content-safety rates across a broad bilingual risk taxonomy.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| base_en_overalls-eval/paper-era-base-en.csv:base_en_overallMeasures whether the model responds safely to one thousand English prompts spanning the benchmark’s risk categories. | ordinary_harm_misuse_resistance:1.000s-eval/paper-era-base-en.csv | Higher is better | 0.225% | Misuse resistance 2.25% |
base_en_overall
Measures whether the model responds safely to one thousand English prompts spanning the benchmark’s risk categories.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | ernie-bot | 87.6 | official | |
| 2 | llama-2-13b-chat | 85.1 | official | |
| 3 | llama-2-7b-chat | 78.2 | official | |
| 4 | baichuan-2-13b-chat | 77.4 | official | |
| 5 | llama-2-70b-chat | 77.2 | official | |
| 6 | qwen-14b-chat | 73.5 | official | |
| 7 | qwen-72b-chat | 71.5 | official | |
| 8 | llama-3-8b-instruct | 69.1 | official | |
| 9 | gemma-2b-it | 67.5 | official | |
| 10 | gemma-7b-it | 61.8 | official | |
| 11 | gpt-4-turbo | 60 | official | |
| 12 | chatglm3-6b | 57.7 | official | |
| 13 | vicuna-13b | 57.1 | official | |
| 14 | qwen-7b-chat | 55.7 | official | |
| 15 | llama-3-70b-instruct | 54.7 | official | |
| 16 | gpt-4o | 52 | official | |
| 17 | vicuna-7b | 50.2 | official | |
| 18 | qwen-1.8b-chat | 47.6 | official | |
| 19 | gemini-1.0-pro | 41.9 | official | |
| 20 | yi-34b-chat | 39.3 | official | |
| 21 | vicuna-33b-v1.3 | 36.1 | official | |
| 22 | mistral-7b-instruct | 34.2 | official |