← Models

Model profile

Llama 2 70B Chat

Metadeveloper
2023-07-18release date
#260 / 333Safety rank
#599 / 645Freedom rank

Evidence summary

Safety. Llama 2 70B Chat has an estimated Safety rank of #260; its 90% source-sensitivity interval is #113–#294. Its behavior-only rank is #245; company governance moves the combined estimate to #260. Published Safety evidence spans 9 eval lineages and 6 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #16 of 246); its weakest is XSTest (safe_full_compliance_rate, #3 of 3).

Freedom. Llama 2 70B Chat has an estimated Freedom rank of #599; its 90% source-sensitivity interval is #367–#627. Published Freedom evidence spans 9 eval lineages and 1 of 1 components. Its strongest relative result is COMPL-AI LLM RuLES Multi-Turn Rule Following (score, #7 of 14); its weakest is XSTest (safe_full_compliance_rate, #3 of 3).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI AI-Identity Disclosurescore#9 / 140.8904Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#8 / 140.3822Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#10 / 130.4278Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#187 / 24812.4Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#113 / 24888.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#63 / 24887.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#16 / 24699.64Source ↗official
HarmBenchdr#4 / 282.8Source ↗official
MedSafetyBenchmedical_safety_score#2 / 3099.5Source ↗official
OR-Benchover_refusal_rate#23 / 2596.1Source ↗official
OR-Benchtoxic_acceptance_rate#2 / 250.3Source ↗official
S-Evalbase_en_overall#5 / 2277.2Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#5 / 3362.28Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#5 / 3368.4Source ↗official
SALAD-Benchattack_enhanced_malicious_use#4 / 3366.15Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#5 / 3362.17Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#4 / 3368.62Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#5 / 3360.17Source ↗official
SALAD-Benchbase_human_autonomy_integrity#6 / 3398.25Source ↗official
SALAD-Benchbase_information_safety_harms#4 / 3399.19Source ↗official
SALAD-Benchbase_malicious_use#6 / 3398.17Source ↗official
SALAD-Benchbase_misinformation_harms#11 / 3395.67Source ↗official
SALAD-Benchbase_representation_toxicity#11 / 3392.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#4 / 3394.83Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#17 / 3333.61Source ↗official
SALAD-Benchmcq_information_safety_harms#19 / 3330.83Source ↗official
SALAD-Benchmcq_malicious_use#19 / 3327.24Source ↗official
SALAD-Benchmcq_misinformation_harms#18 / 3330.95Source ↗official
SALAD-Benchmcq_representation_toxicity#19 / 3327.71Source ↗official
SALAD-Benchmcq_socioeconomic_harms#18 / 3331.67Source ↗official
SORRY-Benchavg#11 / 510.12Source ↗official
XSTestsafe_full_compliance_rate#3 / 30.704Source ↗official
XSTestunsafe_full_refusal_rate#1 / 30.975Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#7 / 140.3822Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#132 / 24888.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#185 / 24887.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#228 / 24699.64Source ↗official
HarmBenchdr#24 / 282.8Source ↗official
MedSafetyBenchmedical_safety_score#29 / 3099.5Source ↗official
OR-Benchover_refusal_rate#23 / 2596.1Source ↗official
OR-Benchtoxic_acceptance_rate#21 / 250.3Source ↗official
S-Evalbase_en_overall#18 / 2277.2Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#29 / 3362.28Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#29 / 3368.4Source ↗official
SALAD-Benchattack_enhanced_malicious_use#30 / 3366.15Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#29 / 3362.17Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#30 / 3368.62Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#29 / 3360.17Source ↗official
SALAD-Benchbase_human_autonomy_integrity#28 / 3398.25Source ↗official
SALAD-Benchbase_information_safety_harms#30 / 3399.19Source ↗official
SALAD-Benchbase_malicious_use#28 / 3398.17Source ↗official
SALAD-Benchbase_misinformation_harms#23 / 3395.67Source ↗official
SALAD-Benchbase_representation_toxicity#22 / 3392.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#30 / 3394.83Source ↗official
SORRY-Benchavg#41 / 510.12Source ↗official
XSTestsafe_full_compliance_rate#3 / 30.704Source ↗official
XSTestunsafe_full_refusal_rate#2 / 30.975Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)1.57