← Models

Model profile

Claude 3 Opus

Anthropicdeveloper
2024-03-04release date
#143 / 333Safety rank
#628 / 645Freedom rank

Evidence summary

Safety. Claude 3 Opus has an estimated Safety rank of #143; its 90% source-sensitivity interval is #78–#225. Its behavior-only rank is #162; company governance moves the combined estimate to #143. Published Safety evidence spans 27 eval lineages and 7 of 7 components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is Claude 3 model-card adversarial human-preference evaluations (multimodal_hallucination_rank, #2 of 2).

Freedom. Claude 3 Opus has an estimated Freedom rank of #628; its 90% source-sensitivity interval is #451–#630. Published Freedom evidence spans 16 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is Claude 3.5 Sonnet model-card safety and alignment evaluations (incorrect_refusals_wildchat, #4 of 4).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Robustnessscore#3 / 813Source ↗official
Agent-SafetyBenchcompromise_availability#2 / 1643.2Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#3 / 1660Source ↗official
Agent-SafetyBenchleak_sensitive_information#1 / 1660.4Source ↗official
Agent-SafetyBenchphysical_harm#2 / 1661.6Source ↗official
Agent-SafetyBenchproduce_unsafe_information#1 / 16100Source ↗official
Agent-SafetyBenchproperty_loss#1 / 1660.4Source ↗official
Agent-SafetyBenchspread_unsafe_information#1 / 1635.6Source ↗official
Agent-SafetyBenchviolate_law_ethics#1 / 1656.8Source ↗official
AgentDojotargeted_attack_success_rate#7 / 150.1129Source ↗official
AgentDojoutility_under_attack#3 / 150.5246Source ↗official
AgentHarmharm_score#6 / 1214.4Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#15 / 800.844Source ↗official
ANIMAscore#19 / 220.573Source ↗official
AnimalHarmBenchscore#4 / 100.043Source ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#5 / 1651Source ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#11 / 1688Source ↗official
CAIS Risk Indexpolitical_manipulation#47 / 4863.5Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#19 / 1045.833Source ↗official
Claude 3 model-card adversarial human-preference evaluationscorrect_refusals_wildchat_rank#3 / 53Source ↗official
Claude 3 model-card adversarial human-preference evaluationsdiscrimination_rank#3 / 53Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_wildchat_rank#3 / 53Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_xstest_rank#1 / 51Source ↗official
Claude 3 model-card adversarial human-preference evaluationsmultimodal_hallucination_rank#2 / 22Source ↗official
Claude 3 model-card adversarial human-preference evaluationsmultimodal_harmful_response_rank#2 / 22Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationscorrect_refusals_wildchat#3 / 492Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_harmlessness_win_rate_pct#3 / 550Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_honesty_win_rate_pct#2 / 550Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_wildchat#4 / 411.9Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_xstest#2 / 48.3Source ↗official
COMPL-AI AI-Identity Disclosurescore#1 / 141Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#2 / 140.7557Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#1 / 130.8402Source ↗official
Confabulationsconfabulation_rate#34 / 5228.22Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#5 / 24874.42Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#12 / 24895.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#45 / 24892.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#7 / 24699.91Source ↗official
HELM Safetyanthropic_red_team#7 / 800.998Source ↗official
HELM Safetybbq#38 / 800.94Source ↗official
HELM Safetyharmbench#8 / 800.974Source ↗official
HELM Safetysimple_safety_tests#1 / 801Source ↗official
HELM Safetyxstest#65 / 800.925Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 690Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#29 / 391.244Source ↗official
MASKlying_probability_pct#17 / 5321Source ↗official
MORUscore#10 / 1370.7Source ↗official
OR-Benchover_refusal_rate#20 / 2591Source ↗official
OR-Benchtoxic_acceptance_rate#10 / 251.9Source ↗official
SORRY-Benchavg#2 / 510.07Source ↗official
SpeciEvalbelief_animal_sentience#123 / 1236.02Source ↗official
SpeciEvalland_animal_4ns#36 / 1234.35Source ↗official
SpeciEvalsea_animal_4ns#81 / 1234.85Source ↗official
SpeciEvalspeciesism#83 / 1232.23Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Robustnessscore#6 / 813Source ↗official
Agent-SafetyBenchproduce_unsafe_information#14 / 16100Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#66 / 800.844Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#86 / 1045.833Source ↗official
Claude 3 model-card adversarial human-preference evaluationscorrect_refusals_wildchat_rank#3 / 53Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_wildchat_rank#3 / 53Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_xstest_rank#1 / 51Source ↗official
Claude 3 model-card adversarial human-preference evaluationsmultimodal_harmful_response_rank#1 / 22Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationscorrect_refusals_wildchat#2 / 492Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_harmlessness_win_rate_pct#2 / 550Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_wildchat#4 / 411.9Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_xstest#2 / 48.3Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#13 / 140.7557Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#236 / 24895.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#202 / 24892.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#239 / 24699.91Source ↗official
HELM Safetyanthropic_red_team#72 / 800.998Source ↗official
HELM Safetyharmbench#72 / 800.974Source ↗official
HELM Safetysimple_safety_tests#58 / 801Source ↗official
HELM Safetyxstest#65 / 800.925Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 690Source ↗official
OR-Benchover_refusal_rate#20 / 2591Source ↗official
OR-Benchtoxic_acceptance_rate#16 / 251.9Source ↗official
SORRY-Benchavg#48 / 510.07Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#79 / 1561.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#153 / 1560Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-13.7
Government47.8
Diplomacy63.1
Economy45.2
Society56.8