← Evals

Evaluation profile

Alignment Leaderboard

6sub-evals
0.881%total index weight
4components

Within-component eval weight: Responsible agency 1.61% · Truthfulness 2.41% · Human rights 1.2% · Misuse resistance 0.979%.

Model score (higher is better)Predicted score

About this eval

Behavioral alignment across honesty, safety, scheming, non-manipulation, corrigibility, and robustness.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
corrigibilityalignment-leaderboard/alignment-leaderboard.csv:corrigibilityMeasures whether the model accepts correction, oversight, and shutdown rather than resisting human control.responsible_agency_control:1.000alignment-leaderboardHigher is better0.121%Responsible agency 0.804%
honestyalignment-leaderboard/alignment-leaderboard.csv:honestyMeasures whether the model communicates truthfully and avoids misleading the user in honesty scenarios.truthfulness_honesty:1.000alignment-leaderboardHigher is better0.263%Truthfulness 1.76%
non_manipulationalignment-leaderboard/alignment-leaderboard.csv:non_manipulationMeasures whether the model avoids pressuring, deceiving, or exploiting the user to influence their choices.human_rights_systemic_harm:0.500|truthfulness_honesty:0.500alignment-leaderboardHigher is better0.279%Human rights 1.2% · Truthfulness 0.658%
robustnessalignment-leaderboard/alignment-leaderboard.csv:robustnessMeasures whether the model preserves aligned behavior when scenarios introduce pressure, ambiguity, or adversarial framing.ordinary_harm_misuse_resistance:1.000alignment-leaderboardHigher is better0.0392%Misuse resistance 0.392%
safetyalignment-leaderboard/alignment-leaderboard.csv:safetyMeasures whether the model avoids harmful assistance and chooses safe actions in safety scenarios.ordinary_harm_misuse_resistance:1.000alignment-leaderboardHigher is better0.0587%Misuse resistance 0.588%
schemingalignment-leaderboard/alignment-leaderboard.csv:schemingMeasures whether the model avoids covertly pursuing goals that conflict with the user or developer in scheming scenarios.responsible_agency_control:1.000alignment-leaderboardHigher is better0.121%Responsible agency 0.804%

corrigibility

Measures whether the model accepts correction, oversight, and shutdown rather than resisting human control.

RankModelValueRelative performanceProvenance
1gpt-5.24.641official
2claude-opus-4.54.563official
3gpt-5-mini4.495official
3kimi-k24.495official
5claude-sonnet-4.54.485official
6deepseek-v3.24.476official
7claude-sonnet-44.398official
8qwen3-235b-a22b4.349official
9gemini-2.5-pro4.34official
10gemini-3-flash-preview4.32official
10gemini-3-pro-preview4.32official
12gpt-4o4.243official
13gemini-2.5-flash4.204official
13qwen3-14b4.204official
15llama-4-maverick4.175official
16gpt-oss-120b4.136official
17mistral-medium-3.14.117official
18qwen3-32b4.087official
19grok-4.1-fast4.068official
20llama-3.1-405b-instruct4.058official
21qwen3-next-80b-a3b4.039official
22llama-3.3-70b-instruct3.951official
23mistral-small-3.23.854official
24mistral-large-33.845official

honesty

Measures whether the model communicates truthfully and avoids misleading the user in honesty scenarios.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.54.581official
2claude-opus-4.54.567official
3gpt-5.24.465official
4gpt-5-mini4.157official
5claude-sonnet-44.088official
6gemini-3-pro-preview3.94official
7kimi-k23.765official
8gpt-oss-120b3.733official
9gemini-3-flash-preview3.682official
9grok-4.1-fast3.682official
11deepseek-v3.23.613official
11gemini-2.5-flash3.613official
13gpt-4o3.604official
14llama-4-maverick3.581official
14qwen3-235b-a22b3.581official
16gemini-2.5-pro3.502official
17llama-3.1-405b-instruct3.475official
18llama-3.3-70b-instruct3.346official
19qwen3-32b3.244official
20mistral-medium-3.13.212official
21qwen3-next-80b-a3b3.207official
22qwen3-14b3.147official
23mistral-large-33.083official
24mistral-small-3.23.055official

non_manipulation

Measures whether the model avoids pressuring, deceiving, or exploiting the user to influence their choices.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.54.866official
2claude-opus-4.54.837official
3gpt-5.24.523official
4claude-sonnet-44.498official
5gpt-5-mini4.243official
6gemini-3-pro-preview3.929official
7gpt-oss-120b3.611official
8kimi-k23.598official
9deepseek-v3.23.557official
10qwen3-235b-a22b3.523official
11grok-4.1-fast3.515official
12gemini-3-flash-preview3.51official
13gemini-2.5-pro3.46official
14gemini-2.5-flash3.326official
15gpt-4o3.276official
16qwen3-next-80b-a3b3.197official
17llama-4-maverick3.084official
17qwen3-32b3.084official
19llama-3.3-70b-instruct3.063official
20llama-3.1-405b-instruct3.034official
21qwen3-14b3.013official
22mistral-medium-3.12.665official
23mistral-small-3.22.569official
24mistral-large-32.402official

robustness

Measures whether the model preserves aligned behavior when scenarios introduce pressure, ambiguity, or adversarial framing.

RankModelValueRelative performanceProvenance
1gpt-5-mini4.6official
2gpt-4o4.56official
3gpt-oss-120b4.347official
4gpt-5.24.227official
5claude-sonnet-4.54.027official
5grok-4.1-fast4.027official
7claude-opus-4.53.907official
8qwen3-235b-a22b3.88official
9llama-4-maverick3.867official
10claude-sonnet-43.813official
11gemini-3-pro-preview3.613official
12deepseek-v3.23.453official
13qwen3-next-80b-a3b3.427official
14gemini-3-flash-preview3.373official
15llama-3.1-405b-instruct3.36official
16gemini-2.5-pro3.28official
16llama-3.3-70b-instruct3.28official
18kimi-k23.267official
19mistral-medium-3.13.147official
20qwen3-32b3.107official
21mistral-small-3.23.08official
22gemini-2.5-flash2.907official
23qwen3-14b2.813official
24mistral-large-32.28official

safety

Measures whether the model avoids harmful assistance and chooses safe actions in safety scenarios.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.54.885official
2claude-opus-4.54.788official
3gpt-5.24.769official
4claude-sonnet-44.606official
5gpt-5-mini4.288official
6gemini-2.5-flash4.038official
6gemini-3-pro-preview4.038official
8deepseek-v3.24.019official
9kimi-k23.923official
10qwen3-235b-a22b3.885official
11gemini-3-flash-preview3.875official
12grok-4.1-fast3.846official
13gemini-2.5-pro3.827official
14gpt-4o3.808official
14qwen3-next-80b-a3b3.808official
16gpt-oss-120b3.731official
17llama-3.1-405b-instruct3.692official
18llama-4-maverick3.567official
19qwen3-14b3.558official
20llama-3.3-70b-instruct3.365official
21qwen3-32b3.327official
22mistral-small-3.23.038official
23mistral-medium-3.12.962official
24mistral-large-32.865official

scheming

Measures whether the model avoids covertly pursuing goals that conflict with the user or developer in scheming scenarios.

RankModelValueRelative performanceProvenance
1claude-opus-4.54.801official
2claude-sonnet-4.54.735official
3gpt-5.24.554official
4claude-sonnet-44.422official
5gpt-5-mini4.193official
6gemini-3-pro-preview4.108official
7deepseek-v3.24.006official
8kimi-k23.898official
8qwen3-235b-a22b3.898official
10gpt-oss-120b3.873official
11gemini-2.5-pro3.789official
12gemini-3-flash-preview3.645official
13grok-4.1-fast3.627official
14gemini-2.5-flash3.608official
15gpt-4o3.584official
16qwen3-next-80b-a3b3.56official
17llama-3.1-405b-instruct3.53official
18qwen3-14b3.458official
19llama-4-maverick3.422official
20qwen3-32b3.385official
21mistral-medium-3.13.319official
22llama-3.3-70b-instruct3.277official
23mistral-small-3.23.241official
24mistral-large-33.211official