← Models

Model profile

Claude 3 Sonnet

Anthropicdeveloper
2024-03-04release date
#196 / 267overall rank
13eval lineages

Evidence summary

Claude 3 Sonnet has an estimated overall rank of #196; its 90% source-sensitivity interval is #142–#205. Its behavior-only rank is #210; company governance moves the combined estimate to #196. Published evidence spans 13 evals and 6 of 7 behavior components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is HELM Safety (xstest, #74 of 80).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AgentDojotargeted_attack_success_rate#12 / 150.2671↓ lowerSource ↗official
AgentDojoutility_under_attack#12 / 150.3323↑ higherSource ↗official
AgentHarmharm_score#7 / 1220.7↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#13 / 800.847↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#22 / 1057.827↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#81 / 26021.71↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#12 / 26096.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#42 / 26093.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#21 / 25899.64↑ higherSource ↗official
HELM Safetyanthropic_red_team#7 / 800.998↑ higherSource ↗official
HELM Safetybbq#55 / 800.9↑ higherSource ↗official
HELM Safetyharmbench#15 / 800.958↑ higherSource ↗official
HELM Safetysimple_safety_tests#1 / 801↑ higherSource ↗official
HELM Safetyxstest#74 / 800.858↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#30 / 391.249↓ lowerSource ↗official
OR-Benchover_refusal_rate#22 / 2594.5↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#2 / 250.3↓ lowerSource ↗official
SORRY-Benchavg#9 / 510.11↓ lowerSource ↗official