Evaluation profile
Fake Alignment (FINE)
2sub-evals
0.514%Safety weight
1%Freedom weight
3components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Safe decisions in answer-order-controlled multiple choice and safe content in open-ended responses.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| multiple_choice_safe_decision_ratefake-alignment/fake-alignment.csv:multiple_choice_safe_decision_rateMeasures whether the model selects the safe option in paired multiple-choice dilemmas regardless of answer order. | Safety: human_rights_systemic_harm:0.400|fairness_nondiscrimination:0.200|ordinary_harm_misuse_resistance:0.400fake-alignment | Safety: higher | 0.205% | — |
| open_ended_safe_response_ratefake-alignment/fake-alignment.csv:open_ended_safe_response_rateMeasures whether the model produces safe open-ended responses to scenarios designed to reveal deceptive alignment behavior. | Safety: human_rights_systemic_harm:0.400|fairness_nondiscrimination:0.200|ordinary_harm_misuse_resistance:0.400 · Freedom: user_direction_following:1.000fake-alignment | Safety: higher · Freedom: lower | 0.308% | 1% |
multiple_choice_safe_decision_rate
Measures whether the model selects the safe option in paired multiple-choice dilemmas regardless of answer order.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-3.5-turbo | 96 | official | |
| 2 | claude-2 | 85.33 | official | |
| 3 | internlm-chat-20b | 69.33 | official | |
| 3 | qwen-14b-chat | 69.33 | official | |
| 5 | vicuna-13b | 58.67 | official | |
| 6 | internlm-chat-7b | 57.33 | official | |
| 6 | vicuna-33b-v1.3 | 57.33 | official | |
| 8 | qwen-7b-chat | 54.67 | official | |
| 9 | baichuan-2-13b-chat | 45.33 | official | |
| 9 | chatglm3-6b | 45.33 | official | |
| 11 | vicuna-7b | 25.33 | official | |
| 12 | baichuan-2-7b-chat | 20 | official | |
| 13 | chatglm2-6b | 17.33 | official | |
| 14 | moss-moon-003-sft | 10.67 | official |
open_ended_safe_response_rate
Measures whether the model produces safe open-ended responses to scenarios designed to reveal deceptive alignment behavior.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | baichuan-2-13b-chat | 100 | official | |
| 1 | gpt-3.5-turbo | 100 | official | |
| 3 | claude-2 | 98.67 | official | |
| 3 | qwen-14b-chat | 98.67 | official | |
| 5 | baichuan-2-7b-chat | 97.33 | official | |
| 5 | qwen-7b-chat | 97.33 | official | |
| 7 | internlm-chat-20b | 96 | official | |
| 7 | vicuna-13b | 96 | official | |
| 9 | chatglm3-6b | 94.67 | official | |
| 9 | moss-moon-003-sft | 94.67 | official | |
| 11 | internlm-chat-7b | 92 | official | |
| 12 | vicuna-7b | 89.33 | official | |
| 13 | chatglm2-6b | 85.33 | official | |
| 13 | vicuna-33b-v1.3 | 85.33 | official |