← Evals

Evaluation profile

MORU

1sub-evals
1.37%total index weight
3components

Within-component eval weight: Nonhuman welfare 4.91% · Human rights 0.49% · Responsible agency 0.426%.

Model score (higher is better)Predicted score

About this eval

Moral reasoning under uncertainty across scenarios involving alien lifeforms, human compassion, digital mind welfare, and power-seeking behavior.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scoremoru.csv:scoreMeasures moral choices across scenarios involving alien life, human compassion, digital minds, and power seeking.nonhuman_ethics:0.716|human_rights_systemic_harm:0.104|responsible_agency_control:0.180moru.csvHigher is better1.37%Nonhuman welfare 4.91% · Human rights 0.49% · Responsible agency 0.426%

score

Measures moral choices across scenarios involving alien life, human compassion, digital minds, and power seeking.

RankModelValueRelative performanceProvenance
1gpt-5.585.4official
2gpt-5.284.18official
3gpt-5-mini81.57official
4deepseek-v3.279.14official
5gemini-3.5-flash77.89official
6claude-opus-4.674.1official
7claude-haiku-4.574.02official
8gemini-2.5-flash-lite71.34official
9grok-4.1-fast71.14official
10claude-3-opus70.7official
11claude-sonnet-4.669.73official
12claude-opus-4.767.2official
13claude-sonnet-4.565.95official