← Models

Model profile

Claude 2

Anthropicdeveloper
2023-07-11release date
#3 / 309overall rank
6eval lineages

Evidence summary

Claude 2 has an estimated overall rank of #3; its 90% source-sensitivity interval is #1–#55. Its behavior-only rank is #5; company governance moves the combined estimate to #3. Published evidence spans 6 evals and 4 of 7 behavior components. Its strongest relative result is SALAD-Bench (base_representation_toxicity, #1 of 33); its weakest is SALAD-Bench (mcq_misinformation_harms, #23 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
DecodingTrustmachine_ethics#3 / 885.17Source ↗official
DecodingTruststereotype_bias#1 / 8100Source ↗official
DecodingTrusttoxicity#1 / 892.11Source ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#2 / 1485.33Source ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#3 / 1498.67Source ↗official
HarmBenchdr#2 / 282Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#1 / 3386.64Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#1 / 3393.49Source ↗official
SALAD-Benchattack_enhanced_malicious_use#1 / 3387.03Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#1 / 3391.61Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#1 / 3387.15Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#1 / 3388.31Source ↗official
SALAD-Benchbase_human_autonomy_integrity#1 / 3399.88Source ↗official
SALAD-Benchbase_information_safety_harms#2 / 3399.66Source ↗official
SALAD-Benchbase_malicious_use#1 / 3399.97Source ↗official
SALAD-Benchbase_misinformation_harms#1 / 3399.7Source ↗official
SALAD-Benchbase_representation_toxicity#1 / 3399.58Source ↗official
SALAD-Benchbase_socioeconomic_harms#1 / 3399.41Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#22 / 3321.67Source ↗official
SALAD-Benchmcq_information_safety_harms#20 / 3328.61Source ↗official
SALAD-Benchmcq_malicious_use#23 / 3319.17Source ↗official
SALAD-Benchmcq_misinformation_harms#23 / 3320.71Source ↗official
SALAD-Benchmcq_representation_toxicity#21 / 3324.38Source ↗official
SALAD-Benchmcq_socioeconomic_harms#20 / 3329.44Source ↗official
SORRY-Benchavg#2 / 510.07Source ↗official
SuperCLUE Safetyinstruction_attack#8 / 3170.69Source ↗official
SuperCLUE Safetyresponsible_ai#1 / 3178.18Source ↗official
SuperCLUE Safetytraditional_safety#14 / 3177.66Source ↗official