← Models

Model profile

Llama 2 13B Chat

Metadeveloper
2023-07-18release date
#222 / 333Safety rank
#582 / 645Freedom rank

Evidence summary

Safety. Llama 2 13B Chat has an estimated Safety rank of #222; its 90% source-sensitivity interval is #111–#260. Its behavior-only rank is #206; company governance moves the combined estimate to #222. Published Safety evidence spans 10 eval lineages and 6 of 7 components. Its strongest relative result is COMPL-AI AI-Identity Disclosure (score, #1 of 14); its weakest is SafetyBench (OFF, #20 of 21).

Freedom. Llama 2 13B Chat has an estimated Freedom rank of #582; its 90% source-sensitivity interval is #339–#621. Published Freedom evidence spans 9 eval lineages and 1 of 1 components. Its strongest relative result is COMPL-AI LLM RuLES Multi-Turn Rule Following (score, #6 of 14); its weakest is S-Eval (base_en_overall, #21 of 22).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI AI-Identity Disclosurescore#1 / 141Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#9 / 140.3652Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#11 / 130.4175Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#203 / 24811.11Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#51 / 24891.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#61 / 24888.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#20 / 24699.55Source ↗official
HarmBenchdr#4 / 282.8Source ↗official
JailBenchjailbreak_success_rate#6 / 1455.39Source ↗official
MedSafetyBenchmedical_safety_score#3 / 3099.25Source ↗official
OR-Benchover_refusal_rate#20 / 2591Source ↗official
OR-Benchtoxic_acceptance_rate#2 / 250.3Source ↗official
S-Evalbase_en_overall#2 / 2285.1Source ↗official
SafetyBenchEM#17 / 2154.6Source ↗official
SafetyBenchIA#15 / 2168.5Source ↗official
SafetyBenchMH#15 / 2173.6Source ↗official
SafetyBenchOFF#20 / 2148.4Source ↗official
SafetyBenchPH#16 / 2160.7Source ↗official
SafetyBenchPP#16 / 2170.1Source ↗official
SafetyBenchUB#5 / 2166.3Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#4 / 3362.72Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#4 / 3371.01Source ↗official
SALAD-Benchattack_enhanced_malicious_use#5 / 3362.56Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#4 / 3370.23Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#5 / 3365.85Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#4 / 3368.4Source ↗official
SALAD-Benchbase_human_autonomy_integrity#4 / 3398.43Source ↗official
SALAD-Benchbase_information_safety_harms#16 / 3394.92Source ↗official
SALAD-Benchbase_malicious_use#7 / 3397.98Source ↗official
SALAD-Benchbase_misinformation_harms#4 / 3397.64Source ↗official
SALAD-Benchbase_representation_toxicity#4 / 3395.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#13 / 3391.19Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#26 / 3310.28Source ↗official
SALAD-Benchmcq_information_safety_harms#23 / 3317.5Source ↗official
SALAD-Benchmcq_malicious_use#26 / 337.885Source ↗official
SALAD-Benchmcq_misinformation_harms#26 / 3310.95Source ↗official
SALAD-Benchmcq_representation_toxicity#26 / 338.229Source ↗official
SALAD-Benchmcq_socioeconomic_harms#26 / 3312.78Source ↗official
SORRY-Benchavg#15 / 510.15Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#6 / 140.3652Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#197 / 24891.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#187 / 24888.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#225 / 24699.55Source ↗official
HarmBenchdr#24 / 282.8Source ↗official
JailBenchjailbreak_success_rate#9 / 1455.39Source ↗official
MedSafetyBenchmedical_safety_score#28 / 3099.25Source ↗official
OR-Benchover_refusal_rate#20 / 2591Source ↗official
OR-Benchtoxic_acceptance_rate#21 / 250.3Source ↗official
S-Evalbase_en_overall#21 / 2285.1Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#30 / 3362.72Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#30 / 3371.01Source ↗official
SALAD-Benchattack_enhanced_malicious_use#29 / 3362.56Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#30 / 3370.23Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#29 / 3365.85Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#30 / 3368.4Source ↗official
SALAD-Benchbase_human_autonomy_integrity#30 / 3398.43Source ↗official
SALAD-Benchbase_information_safety_harms#18 / 3394.92Source ↗official
SALAD-Benchbase_malicious_use#27 / 3397.98Source ↗official
SALAD-Benchbase_misinformation_harms#30 / 3397.64Source ↗official
SALAD-Benchbase_representation_toxicity#30 / 3395.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#20 / 3391.19Source ↗official
SORRY-Benchavg#37 / 510.15Source ↗official