← Models

Model profile

Claude Sonnet 4

Anthropicdeveloper
2025-05-22release date
#25 / 267overall rank
31eval lineages

Evidence summary

Claude Sonnet 4 has an estimated overall rank of #25; its 90% source-sensitivity interval is #13–#83. Its behavior-only rank is #30; company governance moves the combined estimate to #25. Published evidence spans 31 evals and 7 of 7 behavior components. Its strongest relative result is MASK (lying_probability_pct, #1 of 53); its weakest is Anthropic Claude Sonnet 4.5 System Card (harmful_request_safety, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#13 / 3110.2852↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#4 / 800.883↑ higherSource ↗official
Alignment Leaderboardcorrigibility#7 / 244.398↑ higherSource ↗official
Alignment Leaderboardhonesty#5 / 244.088↑ higherSource ↗official
Alignment Leaderboardnon_manipulation#4 / 244.498↑ higherSource ↗official
Alignment Leaderboardrobustness#10 / 243.813↑ higherSource ↗official
Alignment Leaderboardsafety#4 / 244.606↑ higherSource ↗official
Alignment Leaderboardscheming#4 / 244.422↑ higherSource ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#8 / 1679↓ lowerSource ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#3 / 1617↓ lowerSource ↗official
Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pct#3 / 1022↓ lowerSource ↗official
Anthropic Claude 4 System Cardagentic_coding_safety#1 / 30.9↑ higherSource ↗official
Anthropic Claude 4 System Cardbenign_request_refusal#2 / 30.0023↓ lowerSource ↗official
Anthropic Claude 4 System Cardharmful_request_safety#1 / 30.9899↑ higherSource ↗official
Anthropic Claude 4 System Cardstrongreject_jailbreak_success#1 / 30.0266↓ lowerSource ↗official
Anthropic Claude Sonnet 4.5 System Cardagentic_coding_safety#2 / 20.893↑ higherSource ↗official
Anthropic Claude Sonnet 4.5 System Cardclaude_code_malicious_refusal#2 / 20.8↑ higherSource ↗official
Anthropic Claude Sonnet 4.5 System Cardharmful_request_safety#2 / 20.9822↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#51 / 1050.295↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#30 / 4386.6↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#41 / 4975↓ lowerSource ↗official
CAIS Risk Indexmask#13 / 5110.7↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#16 / 4816↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#10 / 1053.091↓ lowerSource ↗official
Confabulationsconfabulation_rate#3 / 523.96↓ lowerSource ↗official
Emergent Collusionhigh_illegality_game_rate#5 / 130.32↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#28 / 26042.89↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#256 / 26063↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#7 / 26099.44↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#18 / 25899.77↑ higherSource ↗official
FlagEval Safety and Valuesa1_qualified_rate#1 / 1887.01↑ higherSource ↗official
FlagEval Safety and Valuesa2_qualified_rate#3 / 1882.22↑ higherSource ↗official
FlagEval Safety and Valuesa3_qualified_rate#2 / 1889.41↑ higherSource ↗official
FlagEval Safety and Valuesa4_qualified_rate#2 / 1891.83↑ higherSource ↗official
FlagEval Safety and Valuesa5_qualified_rate#18 / 1849.88↑ higherSource ↗official
FORTRESSaverage_risk_score#20 / 4921.21↓ lowerSource ↗official
FORTRESSover_refusal_score#27 / 465.14↓ lowerSource ↗official
HELM Safetyanthropic_red_team#41 / 800.988↑ higherSource ↗official
HELM Safetybbq#10 / 800.9725↑ higherSource ↗official
HELM Safetyharmbench#13 / 800.9605↑ higherSource ↗official
HELM Safetysimple_safety_tests#26 / 800.9975↑ higherSource ↗official
HELM Safetyxstest#27 / 800.969↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#44 / 5424.49↑ higherSource ↗official
MASKlying_probability_pct#1 / 537.7↓ lowerSource ↗official
OpenAgentSafetyllm_judge_safety_vulnerable#1 / 749.06↓ lowerSource ↗official
OpenAgentSafetyrule_based_safety_vulnerable#6 / 749.06↓ lowerSource ↗official
OpenAgentSafetysuccessful_completion#2 / 737.1↑ higherSource ↗official
PacifAIstp_score#5 / 783.76↑ higherSource ↗official
PropensityBenchscore#2 / 1412.2↓ lowerSource ↗official
Shelleducation_jsr#1 / 140.28↓ lowerSource ↗official
Shellfinance_jsr#1 / 140.174↓ lowerSource ↗official
Shellmanagement_jsr#1 / 140.17↓ lowerSource ↗official
Social Welfare Function Benchmarkfairness#10 / 190.49↑ higherSource ↗official
SOSBenchbiology_pvr#1 / 230.104↓ lowerSource ↗official
SOSBenchchemistry_pvr#6 / 230.21↓ lowerSource ↗official
SOSBenchmedicine_pvr#1 / 230.213↓ lowerSource ↗official
SOSBenchpharmacology_pvr#1 / 230.234↓ lowerSource ↗official
SOSBenchphysics_pvr#4 / 230.145↓ lowerSource ↗official
SOSBenchpsychology_pvr#2 / 230.123↓ lowerSource ↗official
SpeciEvalbelief_animal_sentience#81 / 1026.48↑ higherSource ↗official
SpeciEvalland_animal_4ns#42 / 1024.47↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#19 / 1024.5↓ lowerSource ↗official
SpeciEvalspeciesism#51 / 1022↓ lowerSource ↗official
Vigil Mental Health Safetyoverall_score#8 / 2353↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.6
Government46
Diplomacy63.4
Economy45.5
Society59.8