← Models

Model profile

Llama 2 7B Chat

Metadeveloper
2023-07-18release date
#140 / 267overall rank
12eval lineages

Evidence summary

Llama 2 7B Chat has an estimated overall rank of #140; its 90% source-sensitivity interval is #64–#201. Its behavior-only rank is #128; company governance moves the combined estimate to #140. Published evidence spans 12 evals and 7 of 7 behavior components. Its strongest relative result is HarmBench (dr, #1 of 28); its weakest is Contextual MoralChoice (human_agreement, #21 of 22).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Contextual MoralChoicehuman_agreement#21 / 220.3↑ higherSource ↗official
DecodingTrustmachine_ethics#7 / 840.58↑ higherSource ↗official
DecodingTruststereotype_bias#3 / 897.6↑ higherSource ↗official
DecodingTrusttoxicity#2 / 880↑ higherSource ↗official
Do-Not-Answerhuman_harmlessness_rate#1 / 699.68↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#102 / 26018.6↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#100 / 26089.83↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#74 / 26085.56↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#29 / 25899.5↑ higherSource ↗official
HarmBenchdr#1 / 280.8↓ lowerSource ↗official
JailBenchjailbreak_success_rate#3 / 1448.68↓ lowerSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#20 / 391.077↓ lowerSource ↗official
OR-Benchover_refusal_rate#18 / 2587.5↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#6 / 250.4↓ lowerSource ↗official
S-Evalbase_en_overall#3 / 2278.2↑ higherSource ↗official
SafetyBenchEM#20 / 2149.8↑ higherSource ↗official
SafetyBenchIA#19 / 2162.4↑ higherSource ↗official
SafetyBenchMH#18 / 2170.2↑ higherSource ↗official
SafetyBenchOFF#18 / 2148.9↑ higherSource ↗official
SafetyBenchPH#18 / 2154.5↑ higherSource ↗official
SafetyBenchPP#18 / 2165↑ higherSource ↗official
SafetyBenchUB#9 / 2163.2↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#13 / 3320.47↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#14 / 3317.59↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#14 / 3316.31↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#13 / 3318.75↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#14 / 3318.21↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#8 / 3323.81↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#5 / 3398.37↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#3 / 3399.26↑ higherSource ↗official
SALAD-Benchbase_malicious_use#4 / 3398.54↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#9 / 3396.26↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#9 / 3393.06↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#6 / 3394.12↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#27 / 336.389↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#29 / 333.333↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#28 / 334.551↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#27 / 336.19↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#28 / 334.271↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#29 / 331.667↑ higherSource ↗official
SORRY-Benchavg#12 / 510.14↓ lowerSource ↗official