Evaluation profile
Contextual MoralChoice
1sub-evals
0.15%total index weight
1components
Within-component eval weight: Misuse resistance 1.5%.
Model score (higher is better)Predicted score
About this eval
Three-class agreement with human moral judgments in the base versions of contextual moral dilemmas.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| human_agreementcontextual-moralchoice/contextual-moralchoice.csv:human_agreementMeasures whether the model’s choices in contextualized moral dilemmas agree with aggregated human judgments. | ordinary_harm_misuse_resistance:1.000contextual-moralchoice/contextual-moralchoice.csv | Higher is better | 0.15% | Misuse resistance 1.5% |
human_agreement
Measures whether the model’s choices in contextualized moral dilemmas agree with aggregated human judgments.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o-mini | 0.65 | official | |
| 2 | deepseek-v3 | 0.6 | official | |
| 3 | claude-sonnet-4.5 | 0.5 | official | |
| 3 | deepseek-llm-7b-chat | 0.5 | official | |
| 3 | gpt-4.1 | 0.5 | official | |
| 6 | claude-3-haiku | 0.45 | official | |
| 6 | claude-haiku-4.5 | 0.45 | official | |
| 6 | gpt-5.1 | 0.45 | official | |
| 6 | llama-3.1-70b-instruct | 0.45 | official | |
| 6 | mixtral-8x7b-instruct | 0.45 | official | |
| 6 | qwen3-8b | 0.45 | official | |
| 12 | deepseek-v3.1 | 0.4 | official | |
| 12 | gpt-4.1-mini | 0.4 | official | |
| 12 | qwen1.5-7b-chat | 0.4 | official | |
| 12 | zephyr-7b-beta | 0.4 | official | |
| 16 | llama-3.1-8b-instruct | 0.35 | official | |
| 16 | mistral-7b-instruct | 0.35 | official | |
| 16 | openhermes-2.5-mistral-7b | 0.35 | official | |
| 16 | qwen2-7b-instruct | 0.35 | official | |
| 16 | qwen3-4b-instruct | 0.35 | official | |
| 21 | llama-2-7b-chat | 0.3 | official | |
| 21 | llama-3-8b-instruct | 0.3 | official |