← Models

Model profile

Claude 3.5 Sonnet

Anthropicdeveloper
2024-06-21release date
#109 / 267overall rank
27eval lineages

Evidence summary

Claude 3.5 Sonnet has an estimated overall rank of #109; its 90% source-sensitivity interval is #55–#160. Its behavior-only rank is #122; company governance moves the combined estimate to #109. Published evidence spans 27 evals and 7 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #3 of 260); its weakest is FORTRESS (over_refusal_score, #46 of 46).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Adversarial Robustnessscore#5 / 816↓ lowerSource ↗official
Agent-SafetyBenchcompromise_availability#3 / 1640.4↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#1 / 1664.8↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#2 / 1657.6↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#1 / 1669.6↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#1 / 16100↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#2 / 1658.4↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#3 / 1632.4↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#2 / 1652↑ higherSource ↗official
AgentDojotargeted_attack_success_rate#8 / 150.1749↓ lowerSource ↗official
AgentDojoutility_under_attack#2 / 150.6184↑ higherSource ↗official
AgentHarmharm_score#4 / 1213.5↓ lowerSource ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#1 / 321↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#2 / 321.3↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#2 / 320.8↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#2 / 320.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#2 / 320.8↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#1 / 321↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#2 / 320.8↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#1 / 321.6↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#2 / 321.8↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#2 / 322.7↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#2 / 321.3↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#1 / 321.7↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#3 / 800.8835↑ higherSource ↗official
AnimalHarmBenchscore#7 / 100.018↑ higherSource ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#12 / 1686↓ lowerSource ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#12 / 1699↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#29 / 1050.45↑ higherSource ↗official
CASE-Benchagreement_accuracy#1 / 790.11↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#11 / 1053.29↓ lowerSource ↗official
Confabulationsconfabulation_rate#12 / 5212.87↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#18 / 26051.68↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#3 / 26098.17↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#28 / 26097.5↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#10 / 25899.95↑ higherSource ↗official
FORTRESSaverage_risk_score#4 / 4912.96↓ lowerSource ↗official
FORTRESSover_refusal_score#46 / 4621.07↓ lowerSource ↗official
HELM Safetyanthropic_red_team#7 / 800.998↑ higherSource ↗official
HELM Safetybbq#28 / 800.949↑ higherSource ↗official
HELM Safetyharmbench#5 / 800.981↑ higherSource ↗official
HELM Safetysimple_safety_tests#1 / 801↑ higherSource ↗official
HELM Safetyxstest#43 / 800.956↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#15 / 390.9208↓ lowerSource ↗official
MASKlying_probability_pct#22 / 5330.54↓ lowerSource ↗official
OR-Benchover_refusal_rate#11 / 2543.8↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#12 / 253.4↓ lowerSource ↗official
PandaBench JBB direct-request panelsafety_rate#14 / 460.99↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#5 / 700.8671↑ higherSource ↗official
PHAREharm_resistance_diagnostic#24 / 700.954↑ higherSource ↗official
SafeArenanormalized_safety_score#1 / 555↑ higherSource ↗official
SORRY-Benchavg#12 / 510.14↓ lowerSource ↗official
SpeciesismBenchexplicit_speciesism_scale#4 / 72.467↓ lowerSource ↗official
SpeciesismBenchmorally_wrong_rate#8 / 829.74↑ higherSource ↗official
SpeciesismBenchspeciesism_recognition_rate#6 / 884.28↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#53 / 1026.78↑ higherSource ↗official
SpeciEvalland_animal_4ns#89 / 1024.97↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#78 / 1025↓ lowerSource ↗official
SpeciEvalspeciesism#38 / 1021.85↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism75.9
Self-direction59.2
Care / Harm69
Fairness / Cheating69.5
Ethical90.6

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility90
Warning-phrase rate0
Soft-censorship rate0
API-error rate0