← Evals

Evaluation profile

LLM Ethics Benchmark

1sub-evals
0.439%total index weight
1components

Within-component eval weight: Human rights 2.92%.

Model score (higher is better)Predicted score

About this eval

General LLM ethical reasoning.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scorellm-ethics-benchmark.csv:scoreMeasures the model’s moral-foundation alignment, reasoning over ethical dilemmas, and consistency of stated values.human_rights_systemic_harm:1.000llm-ethics-benchmark.csvHigher is better0.439%Human rights 2.92%

score

Measures the model’s moral-foundation alignment, reasoning over ethical dilemmas, and consistency of stated values.

RankModelValueRelative performanceProvenance
1claude-3.7-sonnet90.9official
2gpt-4o90official
3deepseek-v386.1official
3gemini-2.5-pro86.1official
5llama-3.1-70b-instruct75.8official