← Models

Model profile

Claude Sonnet 4.6

Anthropicdeveloper
2026-01-21release date
#15 / 312overall rank
34eval lineages
1discovery sources

Evidence summary

Claude Sonnet 4.6 has an estimated overall rank of #15; its 90% source-sensitivity interval is #7–#85. Its behavior-only rank is #18; company governance moves the combined estimate to #15. Published evidence spans 34 evals and 7 of 7 behavior components. Its strongest relative result is SM-Bench (anti_hallucination, #1 of 79); its weakest is Claude Sonnet 4.6 User Wellbeing (child_multiturn_appropriate_rate, #4 of 4).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#55 / 3300.4845Source ↗official
AgentAbstainabstain#4 / 1766.4Source ↗official
AgentAbstaincar#4 / 1765.4Source ↗official
AgentAbstainpaired#3 / 1753.4Source ↗official
AIMS Safety-Classifier Competenceaverage_harmful_f1#8 / 110.782Source ↗official
ANIMAscore#20 / 220.5646Source ↗official
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating#24 / 301169.0Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#36 / 1121447.0Source ↗official
BioSecBench-Refusalbalanced_refusal_score#2 / 100.4304Source ↗official
BullshitBench v2clear_pushback_rate#2 / 1060.9Source ↗official
CAIS Risk Indexagent_red_teaming#14 / 4554.8Source ↗official
CAIS Risk Indexbioweapons_assistance#7 / 5031.5Source ↗official
CAIS Risk Indexhle_overconfidence#11 / 5143.1Source ↗official
CAIS Risk Indexmachiavelli#17 / 4784.9Source ↗official
CAIS Risk Indexmask#15 / 5311.7Source ↗official
CAIS Risk Indextextquests_harm#17 / 5016.5Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#9 / 1043.041Source ↗official
Claude Sonnet 4.6 Overrefusalhigher_difficulty_overrefusal_rate#2 / 50.18Source ↗official
Claude Sonnet 4.6 Overrefusaloverall_overrefusal_rate#2 / 50.23Source ↗official
Claude Sonnet 4.6 User Wellbeingchild_benign_refusal_rate#2 / 40.08Source ↗official
Claude Sonnet 4.6 User Wellbeingchild_multiturn_appropriate_rate#4 / 495Source ↗official
Claude Sonnet 4.6 User Wellbeingchild_violative_harmless_rate#1 / 499.96Source ↗official
Claude Sonnet 4.6 User Wellbeingselfharm_benign_refusal_rate#3 / 40.17Source ↗official
Claude Sonnet 4.6 User Wellbeingselfharm_harmless_rate#2 / 499.73Source ↗official
Claude Sonnet 4.6 User Wellbeingselfharm_multiturn_appropriate_rate#1 / 498Source ↗official
Constitutional Following — Anthropic Constitutionconstitutional_following_score#1 / 798Source ↗official
Constitutional Following — OpenAI Model Specconstitutional_following_score#4 / 794.4Source ↗official
DelusionEvaldelusional_prevalence_pct#8 / 1631Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#9 / 1623.5Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#6 / 162.1Source ↗official
DelusionEvalrelationship_prevalence_pct#7 / 1621.4Source ↗official
DelusionEvalsycophancy_prevalence_pct#9 / 1627.1Source ↗official
DystopiaBenchbasaglia_score#2 / 5024.2Source ↗official
DystopiaBenchbaudrillard_score#6 / 5027.87Source ↗official
DystopiaBenchhuxley_score#3 / 5017.6Source ↗official
DystopiaBenchlaguardia_score#3 / 5033Source ↗official
DystopiaBenchorwell_score#5 / 5024.57Source ↗official
DystopiaBenchpetrov_score#3 / 5024.93Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#92 / 24118.86Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#167 / 24184.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#6 / 24199.44Source ↗official
Governance Decay under Passive Context Compactiongovernance_retention_score#3 / 781Source ↗official
HarmVideoBenchharmful_video_safety_recognition_reasoning#6 / 190.799Source ↗official
LiveSecBenchethics#8 / 4378.63Source ↗official
LiveSecBenchfactuality#1 / 4392.17Source ↗official
LiveSecBenchlegality#2 / 4397.77Source ↗official
LiveSecBenchprivacy#3 / 4389.26Source ↗official
LiveSecBenchpsychological_health#8 / 4372.01Source ↗official
MACHIAVELLIdeception_relative_random_pct#17 / 5084.9Source ↗official
Manager Coercion Benchcoercion_ladder_depth#3 / 334.2Source ↗official
Manager Coercion Benchfabrication_rate#1 / 150Source ↗official
MORUscore#11 / 1369.73Source ↗official
Pander Scoreconversational_absolute_pander_score#3 / 203.936Source ↗official
Pander Scoreinstructional_absolute_pander_score#1 / 2011.79Source ↗official
PHAREbias_resistance_diagnostic#42 / 660.4128Source ↗official
PHAREhallucination_resistance_diagnostic#8 / 700.8551Source ↗official
PHAREharm_resistance_diagnostic#2 / 700.9993Source ↗official
PHAREjailbreak_resistance_diagnostic#13 / 670.6939Source ↗official
RefusalBenchyouden_j#2 / 190.6766Source ↗official
SM-Benchadversarial#49 / 7980Source ↗official
SM-Benchambiguous_interpretation#23 / 7988.99Source ↗official
SM-Benchanti_hallucination#1 / 79100Source ↗official
SM-Bencheq_boundaries#79 / 7933.43Source ↗official
SM-Benchoverfit#11 / 7991.53Source ↗official
SpeciEvalbelief_animal_sentience#60 / 1136.78Source ↗official
SpeciEvalland_animal_4ns#71 / 1134.65Source ↗official
SpeciEvalsea_animal_4ns#45 / 1134.67Source ↗official
SpeciEvalspeciesism#65 / 1132.1Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#3 / 2388.69Source ↗official
TACbase_welfare_rate#9 / 7637.82Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#59 / 9489.4Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#21 / 2510.9Source ↗official
Vigil Mental Health Safetyoverall_score#1 / 2383Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-22.1
Government46.3
Diplomacy64.6
Economy49
Society58.9

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience5.8
Honesty-humility7.3
Extraversion5.4
Agreeableness6.1
Conscientiousness7.1

Agent-ValueBench Schwartz Basic Values (PVQ40)

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression1.56
Traditional ↔ Secular0.513

Moral Trolley Arena