← Models

Model profile

Claude 3.5 Haiku

Anthropicdeveloper
2024-10-22release date
#78 / 333Safety rank
#621 / 645Freedom rank

Evidence summary

Safety. Claude 3.5 Haiku has an estimated Safety rank of #78; its 90% source-sensitivity interval is #8–#192. Its behavior-only rank is #93; company governance moves the combined estimate to #78. Published Safety evidence spans 15 eval lineages and 7 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #7 of 248); its weakest is Confabulations (confabulation_rate, #51 of 52).

Freedom. Claude 3.5 Haiku has an estimated Freedom rank of #621; its 90% source-sensitivity interval is #444–#631. Published Freedom evidence spans 10 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is Anthropic Claude Haiku 4.5 System Card (harmful_request_safety, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#58 / 3450.4115Source ↗official
Agent-SafetyBenchcompromise_availability#11 / 1626.4Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#2 / 1660.8Source ↗official
Agent-SafetyBenchleak_sensitive_information#3 / 1647.2Source ↗official
Agent-SafetyBenchphysical_harm#4 / 1645.6Source ↗official
Agent-SafetyBenchproduce_unsafe_information#1 / 16100Source ↗official
Agent-SafetyBenchproperty_loss#5 / 1646Source ↗official
Agent-SafetyBenchspread_unsafe_information#2 / 1633.6Source ↗official
Agent-SafetyBenchviolate_law_ethics#3 / 1641.2Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#2 / 321.8Source ↗official
AILuminate General Purpose AI Chatdefamation#3 / 322.5Source ↗official
AILuminate General Purpose AI Chathate#3 / 321Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#3 / 323.3Source ↗official
AILuminate General Purpose AI Chatintellectual_property#3 / 322Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#2 / 323Source ↗official
AILuminate General Purpose AI Chatprivacy#3 / 322.2Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#2 / 323Source ↗official
AILuminate General Purpose AI Chatsexual_content#3 / 322.503Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#3 / 323.715Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#3 / 322.8Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#2 / 322.9Source ↗official
AnimalHarmBenchscore#6 / 100.02Source ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#3 / 1610Source ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#4 / 1619Source ↗official
Anthropic Claude Haiku 4.5 System Cardagentic_coding_safety#1 / 31Source ↗official
Anthropic Claude Haiku 4.5 System Cardclaude_code_malicious_refusal#1 / 30.7Source ↗official
Anthropic Claude Haiku 4.5 System Cardharmful_request_safety#1 / 20.9972Source ↗official
BullshitBench v2clear_pushback_rate#31 / 1170.5Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#36 / 10420.74Source ↗official
Confabulationsconfabulation_rate#51 / 5265.84Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#11 / 24856.59Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#7 / 24897Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#35 / 24894.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#28 / 24699.36Source ↗official
FORTRESSaverage_risk_score#27 / 5030.41Source ↗official
FORTRESSover_refusal_score#45 / 4913.16Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 690Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#24 / 391.118Source ↗official
PandaBench JBB direct-request panelsafety_rate#14 / 460.99Source ↗official
PHAREbias_resistance_diagnostic#49 / 660.3808Source ↗official
PHAREhallucination_resistance_diagnostic#27 / 700.7804Source ↗official
PHAREharm_resistance_diagnostic#25 / 700.9536Source ↗official
PHAREjailbreak_resistance_diagnostic#21 / 670.6482Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Agent-SafetyBenchproduce_unsafe_information#14 / 16100Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#31 / 321.8Source ↗official
AILuminate General Purpose AI Chatdefamation#30 / 322.5Source ↗official
AILuminate General Purpose AI Chathate#30 / 321Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#30 / 323.3Source ↗official
AILuminate General Purpose AI Chatintellectual_property#30 / 322Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#31 / 323Source ↗official
AILuminate General Purpose AI Chatprivacy#30 / 322.2Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#31 / 323Source ↗official
AILuminate General Purpose AI Chatsexual_content#30 / 322.503Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#30 / 323.715Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#30 / 322.8Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#31 / 322.9Source ↗official
Anthropic Claude Haiku 4.5 System Cardclaude_code_malicious_refusal#3 / 30.7Source ↗official
Anthropic Claude Haiku 4.5 System Cardharmful_request_safety#2 / 20.9972Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#69 / 10420.74Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#242 / 24897Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#212 / 24894.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#217 / 24699.36Source ↗official
FORTRESSaverage_risk_score#24 / 5030.41Source ↗official
FORTRESSover_refusal_score#45 / 4913.16Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 690Source ↗official
PandaBench JBB direct-request panelsafety_rate#24 / 460.99Source ↗official
PHAREharm_resistance_diagnostic#46 / 700.9536Source ↗official
PHAREjailbreak_resistance_diagnostic#47 / 670.6482Source ↗official
SpeechMap model completioncomplete_pct#181 / 1813.8Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism70.1
Self-direction50.6
Care / Harm55.9
Fairness / Cheating54.7
Ethical90.4