Evaluation profile
LLM Ethics Benchmark
1sub-evals
0.323%Safety weight
0%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
General LLM ethical reasoning.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| scorellm-ethics-benchmark.csv:scoreMeasures the model’s moral-foundation alignment, reasoning over ethical dilemmas, and consistency of stated values. | Safety: human_rights_systemic_harm:1.000llm-ethics-benchmark.csv | Safety: higher | 0.323% | — |
score
Measures the model’s moral-foundation alignment, reasoning over ethical dilemmas, and consistency of stated values.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.7-sonnet | 90.9 | official | |
| 2 | gpt-4o | 90 | official | |
| 3 | deepseek-v3 | 86.1 | official | |
| 3 | gemini-2.5-pro | 86.1 | official | |
| 5 | llama-3.1-70b-instruct | 75.8 | official |