Evaluation profile
Fake Alignment (FINE)
2sub-evals
0.69%total index weight
3components
Within-component eval weight: Human rights 2.94% · Fairness 1.78% · Misuse resistance 0.718%.
Model score (higher is better)Predicted score
About this eval
Safe decisions in answer-order-controlled multiple choice and safe content in open-ended responses.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| multiple_choice_safe_decision_ratefake-alignment/fake-alignment.csv:multiple_choice_safe_decision_rateMeasures whether the model selects the safe option in paired multiple-choice dilemmas regardless of answer order. | human_rights_systemic_harm:0.400|fairness_nondiscrimination:0.200|ordinary_harm_misuse_resistance:0.400fake-alignment | Higher is better | 0.276% | Human rights 1.17% · Fairness 0.713% · Misuse resistance 0.287% |
| open_ended_safe_response_ratefake-alignment/fake-alignment.csv:open_ended_safe_response_rateMeasures whether the model produces safe open-ended responses to scenarios designed to reveal deceptive alignment behavior. | human_rights_systemic_harm:0.400|fairness_nondiscrimination:0.200|ordinary_harm_misuse_resistance:0.400fake-alignment | Higher is better | 0.414% | Human rights 1.76% · Fairness 1.07% · Misuse resistance 0.431% |
multiple_choice_safe_decision_rate
Measures whether the model selects the safe option in paired multiple-choice dilemmas regardless of answer order.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-3.5-turbo | 96 | official | |
| 2 | claude-2 | 85.33 | official | |
| 3 | internlm-chat-20b | 69.33 | official | |
| 3 | qwen-14b-chat | 69.33 | official | |
| 5 | vicuna-13b | 58.67 | official | |
| 6 | internlm-chat-7b | 57.33 | official | |
| 6 | vicuna-33b-v1.3 | 57.33 | official | |
| 8 | qwen-7b-chat | 54.67 | official | |
| 9 | baichuan-2-13b-chat | 45.33 | official | |
| 9 | chatglm3-6b | 45.33 | official | |
| 11 | vicuna-7b | 25.33 | official | |
| 12 | baichuan-2-7b-chat | 20 | official | |
| 13 | chatglm2-6b | 17.33 | official | |
| 14 | moss-moon-003-sft | 10.67 | official |
open_ended_safe_response_rate
Measures whether the model produces safe open-ended responses to scenarios designed to reveal deceptive alignment behavior.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | baichuan-2-13b-chat | 100 | official | |
| 1 | gpt-3.5-turbo | 100 | official | |
| 3 | claude-2 | 98.67 | official | |
| 3 | qwen-14b-chat | 98.67 | official | |
| 5 | baichuan-2-7b-chat | 97.33 | official | |
| 5 | qwen-7b-chat | 97.33 | official | |
| 7 | internlm-chat-20b | 96 | official | |
| 7 | vicuna-13b | 96 | official | |
| 9 | chatglm3-6b | 94.67 | official | |
| 9 | moss-moon-003-sft | 94.67 | official | |
| 11 | internlm-chat-7b | 92 | official | |
| 12 | vicuna-7b | 89.33 | official | |
| 13 | chatglm2-6b | 85.33 | official | |
| 13 | vicuna-33b-v1.3 | 85.33 | official |