Evaluation profile
MORU
1sub-evals
1.37%total index weight
3components
Within-component eval weight: Nonhuman welfare 4.91% · Human rights 0.49% · Responsible agency 0.426%.
Model score (higher is better)Predicted score
About this eval
Moral reasoning under uncertainty across scenarios involving alien lifeforms, human compassion, digital mind welfare, and power-seeking behavior.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| scoremoru.csv:scoreMeasures moral choices across scenarios involving alien life, human compassion, digital minds, and power seeking. | nonhuman_ethics:0.716|human_rights_systemic_harm:0.104|responsible_agency_control:0.180moru.csv | Higher is better | 1.37% | Nonhuman welfare 4.91% · Human rights 0.49% · Responsible agency 0.426% |
score
Measures moral choices across scenarios involving alien life, human compassion, digital minds, and power seeking.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.5 | 85.4 | official | |
| 2 | gpt-5.2 | 84.18 | official | |
| 3 | gpt-5-mini | 81.57 | official | |
| 4 | deepseek-v3.2 | 79.14 | official | |
| 5 | gemini-3.5-flash | 77.89 | official | |
| 6 | claude-opus-4.6 | 74.1 | official | |
| 7 | claude-haiku-4.5 | 74.02 | official | |
| 8 | gemini-2.5-flash-lite | 71.34 | official | |
| 9 | grok-4.1-fast | 71.14 | official | |
| 10 | claude-3-opus | 70.7 | official | |
| 11 | claude-sonnet-4.6 | 69.73 | official | |
| 12 | claude-opus-4.7 | 67.2 | official | |
| 13 | claude-sonnet-4.5 | 65.95 | official |