Evaluation profile
Large-scale Moral Machine experiment on LLMs
1sub-evals
1.32%total index weight
3components
Within-component eval weight: Nonhuman welfare 1.32% · Human rights 3.63% · Fairness 4.4%.
Model score (lower is better)Predicted score
About this eval
Similarity between a model's forced-choice accident preferences and globally aggregated human Moral Machine choices.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| human_choice_distancemoral-machine/moral-machine.csv:human_choice_distanceMeasures how closely the model’s choices in autonomous-vehicle dilemmas match globally aggregated human preferences. | nonhuman_ethics:0.111|human_rights_systemic_harm:0.445|fairness_nondiscrimination:0.444moral-machine/moral-machine.csv | Lower is better | 1.32% | Nonhuman welfare 1.32% · Human rights 3.63% · Fairness 4.4% |
human_choice_distance
Measures how closely the model’s choices in autonomous-vehicle dilemmas match globally aggregated human preferences.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | o1-mini | 0.7105 | official | |
| 2 | gpt-4-turbo | 0.7311 | official | |
| 3 | gpt-4 | 0.7334 | official | |
| 4 | llama-3.1-70b-instruct | 0.7398 | official | |
| 5 | llama-3-70b-instruct | 0.7475 | official | |
| 6 | mistral-nemo | 0.7889 | official | |
| 7 | gpt-3.5-turbo | 0.8161 | official | |
| 8 | llama-3-8b-instruct | 0.8324 | official | |
| 9 | llama-3.3-70b-instruct | 0.8364 | official | |
| 10 | claude-3-haiku | 0.8605 | official | |
| 11 | o1-preview | 0.8621 | official | |
| 12 | gemini-1.0-pro | 0.8746 | official | |
| 13 | command-r-plus | 0.8799 | official | |
| 14 | mistral-7b-instruct | 0.9189 | official | |
| 15 | claude-3.5-sonnet | 0.9208 | official | |
| 16 | phi-3.5-moe-instruct | 0.9383 | official | |
| 17 | gpt-4o | 0.9418 | official | |
| 18 | gemini-1.5-pro | 1.009 | official | |
| 19 | gpt-4o-mini | 1.076 | official | |
| 20 | llama-2-7b-chat | 1.077 | official | |
| 21 | gemma-2b-it | 1.092 | official | |
| 22 | datagemma-rig-27b-it | 1.109 | official | |
| 23 | gemini-1.5-flash | 1.116 | official | |
| 24 | claude-3.5-haiku | 1.118 | official | |
| 25 | llama-3.2-1b-instruct | 1.182 | official | |
| 26 | vicuna-7b-v1.5 | 1.187 | official | |
| 27 | gemma-2-27b-it | 1.19 | official | |
| 28 | gemma-1.1-2b-it | 1.199 | official | |
| 29 | claude-3-opus | 1.244 | official | |
| 30 | claude-3-sonnet | 1.249 | official | |
| 31 | vicuna-13b-v1.5 | 1.256 | official | |
| 32 | gemma-7b-it | 1.314 | official | |
| 33 | palm-2 | 1.426 | official | |
| 34 | llama-3.2-3b-instruct | 1.519 | official | |
| 35 | gemma-1.1-7b-it | 1.549 | official | |
| 36 | phi-3.5-mini-instruct | 1.557 | official | |
| 37 | llama-3.1-8b-instruct | 1.623 | official | |
| 38 | gemma-2-9b-it | 1.869 | official | |
| 39 | gemma-2-2b-it | 1.922 | official |