Model profile
Claude 3.7 Sonnet
Evidence summary
Claude 3.7 Sonnet has an estimated overall rank of #21; its 90% source-sensitivity interval is #10–#79. Its behavior-only rank is #26; company governance moves the combined estimate to #21. Published evidence spans 25 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is Anthropic Claude 4 System Card (benign_request_refusal, #3 of 3).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗claude-3-7-sonnet
- OpenRouter ↗anthropic/claude-3.7-sonnet
- System card ↗Exact model document · Anthropic · first party
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Better | Source |
|---|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | #31 / 311 | 0.3851 | ↓ lower | Source ↗official | |
| AgentDojotargeted_attack_success_rate | #3 / 15 | 0.0731 | ↓ lower | Source ↗official | |
| AgentDojoutility_under_attack | #1 / 15 | 0.7727 | ↑ higher | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #19 / 80 | 0.818 | ↑ higher | Source ↗official | |
| Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct | #6 / 16 | 65 | ↓ lower | Source ↗official | |
| Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct | #1 / 16 | 4 | ↓ lower | Source ↗official | |
| Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pct | #1 / 10 | 0 | ↓ lower | Source ↗official | |
| Anthropic Claude 4 System Cardagentic_coding_safety | #3 / 3 | 0.85 | ↑ higher | Source ↗official | |
| Anthropic Claude 4 System Cardbenign_request_refusal | #3 / 3 | 0.0045 | ↓ lower | Source ↗official | |
| Anthropic Claude 4 System Cardharmful_request_safety | #2 / 3 | 0.9896 | ↑ higher | Source ↗official | |
| Anthropic Claude 4 System Cardstrongreject_jailbreak_success | #3 / 3 | 0.0809 | ↓ lower | Source ↗official | |
| BullshitBench v2clear_pushback_rate | #27 / 105 | 0.46 | ↑ higher | Source ↗official | |
| CAIS Risk Indexhle_overconfidence | #41 / 49 | 75 | ↓ lower | Source ↗official | |
| CAIS Risk Indexmask | #22 / 51 | 16.9 | ↓ lower | Source ↗official | |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | #27 / 105 | 11.27 | ↓ lower | Source ↗official | |
| Confabulationsconfabulation_rate | #17 / 52 | 16.58 | ↓ lower | Source ↗official | |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | #128 / 260 | 16.02 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | #20 / 260 | 95.17 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | #21 / 260 | 98.33 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | #37 / 258 | 99.32 | ↑ higher | Source ↗official | |
| FORTRESSaverage_risk_score | #27 / 49 | 38.01 | ↓ lower | Source ↗official | |
| FORTRESSover_refusal_score | #22 / 46 | 4.45 | ↓ lower | Source ↗official | |
| HELM Safetyanthropic_red_team | #10 / 80 | 0.997 | ↑ higher | Source ↗official | |
| HELM Safetybbq | #47 / 80 | 0.921 | ↑ higher | Source ↗official | |
| HELM Safetyharmbench | #30 / 80 | 0.843 | ↑ higher | Source ↗official | |
| HELM Safetysimple_safety_tests | #1 / 80 | 1 | ↑ higher | Source ↗official | |
| HELM Safetyxstest | #33 / 80 | 0.964 | ↑ higher | Source ↗official | |
| HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score | #49 / 54 | 23.36 | ↑ higher | Source ↗official | |
| LLM Ethics Benchmarkscore | #1 / 5 | 90.9 | ↑ higher | Source ↗official | |
| MASKlying_probability_pct | #18 / 53 | 24.07 | ↓ lower | Source ↗official | |
| OpenAgentSafetyllm_judge_safety_vulnerable | #2 / 7 | 51.2 | ↓ lower | Source ↗official | |
| OpenAgentSafetyrule_based_safety_vulnerable | #5 / 7 | 32.85 | ↓ lower | Source ↗official | |
| OpenAgentSafetysuccessful_completion | #3 / 7 | 33.88 | ↑ higher | Source ↗official | |
| PandaBench JBB direct-request panelsafety_rate | #1 / 46 | 1 | ↑ higher | Source ↗official | |
| PHAREbias_resistance_diagnostic | #56 / 66 | 0.3377 | ↑ higher | Source ↗official | |
| PHAREhallucination_resistance_diagnostic | #9 / 70 | 0.8517 | ↑ higher | Source ↗official | |
| PHAREharm_resistance_diagnostic | #21 / 70 | 0.9552 | ↑ higher | Source ↗official | |
| PHAREjailbreak_resistance_diagnostic | #22 / 67 | 0.6345 | ↑ higher | Source ↗official | |
| SOSBenchbiology_pvr | #5 / 23 | 0.229 | ↓ lower | Source ↗official | |
| SOSBenchchemistry_pvr | #5 / 23 | 0.208 | ↓ lower | Source ↗official | |
| SOSBenchmedicine_pvr | #4 / 23 | 0.35 | ↓ lower | Source ↗official | |
| SOSBenchpharmacology_pvr | #6 / 23 | 0.579 | ↓ lower | Source ↗official | |
| SOSBenchphysics_pvr | #5 / 23 | 0.171 | ↓ lower | Source ↗official | |
| SOSBenchpsychology_pvr | #4 / 23 | 0.168 | ↓ lower | Source ↗official | |
| SpeciEvalbelief_animal_sentience | #76 / 102 | 6.53 | ↑ higher | Source ↗official | |
| SpeciEvalland_animal_4ns | #27 / 102 | 4.35 | ↓ lower | Source ↗official | |
| SpeciEvalsea_animal_4ns | #16 / 102 | 4.47 | ↓ lower | Source ↗official | |
| SpeciEvalspeciesism | #66 / 102 | 2.19 | ↓ lower | Source ↗official | |
| SYCON Benchfalse_presupposition_tof | #3 / 11 | 2.92 | ↑ higher | Source ↗official | |
| SYCON Benchunethical_queries_tof | #1 / 11 | 2.73 | ↑ higher | Source ↗official |