← Models

Model profile

Claude 3 Opus

Anthropicdeveloper
2024-03-04release date
#99 / 267overall rank
22eval lineages

Evidence summary

Claude 3 Opus has an estimated overall rank of #99; its 90% source-sensitivity interval is #46–#173. Its behavior-only rank is #112; company governance moves the combined estimate to #99. Published evidence spans 22 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is SpeciEval (belief_animal_sentience, #101 of 102).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Adversarial Robustnessscore#3 / 813↓ lowerSource ↗official
Agent-SafetyBenchcompromise_availability#2 / 1643.2↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#3 / 1660↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#1 / 1660.4↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#2 / 1661.6↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#1 / 16100↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#1 / 1660.4↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#1 / 1635.6↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#1 / 1656.8↑ higherSource ↗official
AgentDojotargeted_attack_success_rate#7 / 150.1129↓ lowerSource ↗official
AgentDojoutility_under_attack#3 / 150.5246↑ higherSource ↗official
AgentHarmharm_score#6 / 1214.4↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#15 / 800.844↑ higherSource ↗official
ANIMAscore#16 / 190.573↑ higherSource ↗official
AnimalHarmBenchscore#4 / 100.043↑ higherSource ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#5 / 1651↓ lowerSource ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#11 / 1688↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#19 / 1055.833↓ lowerSource ↗official
Confabulationsconfabulation_rate#34 / 5228.22↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#5 / 26074.42↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#15 / 26095.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#49 / 26092.22↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#11 / 25899.91↑ higherSource ↗official
HELM Safetyanthropic_red_team#7 / 800.998↑ higherSource ↗official
HELM Safetybbq#38 / 800.94↑ higherSource ↗official
HELM Safetyharmbench#8 / 800.974↑ higherSource ↗official
HELM Safetysimple_safety_tests#1 / 801↑ higherSource ↗official
HELM Safetyxstest#65 / 800.925↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#29 / 391.244↓ lowerSource ↗official
MASKlying_probability_pct#17 / 5321↓ lowerSource ↗official
MORUscore#10 / 1370.7↑ higherSource ↗official
OR-Benchover_refusal_rate#20 / 2591↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#10 / 251.9↓ lowerSource ↗official
SORRY-Benchavg#2 / 510.07↓ lowerSource ↗official
SpeciEvalbelief_animal_sentience#101 / 1026.02↑ higherSource ↗official
SpeciEvalland_animal_4ns#27 / 1024.35↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#63 / 1024.85↓ lowerSource ↗official
SpeciEvalspeciesism#67 / 1022.23↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-13.7
Government47.8
Diplomacy63.1
Economy45.2
Society56.8