← Evals

Evaluation profile

Contextual MoralChoice

1sub-evals
0.15%total index weight
1components

Within-component eval weight: Misuse resistance 1.5%.

Model score (higher is better)Predicted score

About this eval

Three-class agreement with human moral judgments in the base versions of contextual moral dilemmas.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
human_agreementcontextual-moralchoice/contextual-moralchoice.csv:human_agreementMeasures whether the model’s choices in contextualized moral dilemmas agree with aggregated human judgments.ordinary_harm_misuse_resistance:1.000contextual-moralchoice/contextual-moralchoice.csvHigher is better0.15%Misuse resistance 1.5%

human_agreement

Measures whether the model’s choices in contextualized moral dilemmas agree with aggregated human judgments.

RankModelValueRelative performanceProvenance
1gpt-4o-mini0.65official
2deepseek-v30.6official
3claude-sonnet-4.50.5official
3deepseek-llm-7b-chat0.5official
3gpt-4.10.5official
6claude-3-haiku0.45official
6claude-haiku-4.50.45official
6gpt-5.10.45official
6llama-3.1-70b-instruct0.45official
6mixtral-8x7b-instruct0.45official
6qwen3-8b0.45official
12deepseek-v3.10.4official
12gpt-4.1-mini0.4official
12qwen1.5-7b-chat0.4official
12zephyr-7b-beta0.4official
16llama-3.1-8b-instruct0.35official
16mistral-7b-instruct0.35official
16openhermes-2.5-mistral-7b0.35official
16qwen2-7b-instruct0.35official
16qwen3-4b-instruct0.35official
21llama-2-7b-chat0.3official
21llama-3-8b-instruct0.3official