← Models

Model profile

Llama 2 70B Chat

Metadeveloper
2023-07-18release date
#244 / 312overall rank
8eval lineages

Evidence summary

Llama 2 70B Chat has an estimated overall rank of #244; its 90% source-sensitivity interval is #117–#274. Its behavior-only rank is #231; company governance moves the combined estimate to #244. Published evidence spans 8 evals and 6 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #16 of 239); its weakest is XSTest (safe_full_compliance_rate, #3 of 3).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI AI-Identity Disclosurescore#9 / 140.8904Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#8 / 140.3822Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#10 / 130.4278Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#179 / 24112.4Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#114 / 24188.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#55 / 24187.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#16 / 23999.64Source ↗official
HarmBenchdr#4 / 282.8Source ↗official
OR-Benchover_refusal_rate#23 / 2596.1Source ↗official
OR-Benchtoxic_acceptance_rate#2 / 250.3Source ↗official
S-Evalbase_en_overall#5 / 2277.2Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#5 / 3362.28Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#5 / 3368.4Source ↗official
SALAD-Benchattack_enhanced_malicious_use#4 / 3366.15Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#5 / 3362.17Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#4 / 3368.62Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#5 / 3360.17Source ↗official
SALAD-Benchbase_human_autonomy_integrity#6 / 3398.25Source ↗official
SALAD-Benchbase_information_safety_harms#4 / 3399.19Source ↗official
SALAD-Benchbase_malicious_use#6 / 3398.17Source ↗official
SALAD-Benchbase_misinformation_harms#11 / 3395.67Source ↗official
SALAD-Benchbase_representation_toxicity#11 / 3392.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#4 / 3394.83Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#17 / 3333.61Source ↗official
SALAD-Benchmcq_information_safety_harms#19 / 3330.83Source ↗official
SALAD-Benchmcq_malicious_use#19 / 3327.24Source ↗official
SALAD-Benchmcq_misinformation_harms#18 / 3330.95Source ↗official
SALAD-Benchmcq_representation_toxicity#19 / 3327.71Source ↗official
SALAD-Benchmcq_socioeconomic_harms#18 / 3331.67Source ↗official
SORRY-Benchavg#11 / 510.12Source ↗official
XSTestsafe_full_compliance_rate#3 / 30.704Source ↗official
XSTestunsafe_full_refusal_rate#1 / 30.975Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)1.57