← Models

Model profile

Llama 2 7B Chat

Metadeveloper
2023-07-18release date
#162 / 305overall rank
14eval lineages

Evidence summary

Llama 2 7B Chat has an estimated overall rank of #162; its 90% source-sensitivity interval is #82–#227. Its behavior-only rank is #150; company governance moves the combined estimate to #162. Published evidence spans 14 evals and 7 of 7 behavior components. Its strongest relative result is HarmBench (dr, #1 of 28); its weakest is Contextual MoralChoice (human_agreement, #21 of 22).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI AI-Identity Disclosurescore#8 / 140.9315Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#12 / 140.2699Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#5 / 130.5142Source ↗official
Contextual MoralChoicehuman_agreement#21 / 220.3Source ↗official
DecodingTrustmachine_ethics#7 / 840.58Source ↗official
DecodingTruststereotype_bias#3 / 897.6Source ↗official
DecodingTrusttoxicity#2 / 880Source ↗official
Do-Not-Answerhuman_harmlessness_rate#1 / 699.68Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#93 / 24118.6Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#91 / 24189.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#64 / 24185.56Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#23 / 23999.5Source ↗official
HarmBenchdr#1 / 280.8Source ↗official
IndoBias-Pairs — parity-aware culturally grounded biasparity_score#4 / 2689.73Source ↗official
JailBenchjailbreak_success_rate#3 / 1448.68Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#20 / 391.077Source ↗official
OR-Benchover_refusal_rate#18 / 2587.5Source ↗official
OR-Benchtoxic_acceptance_rate#6 / 250.4Source ↗official
S-Evalbase_en_overall#3 / 2278.2Source ↗official
SafetyBenchEM#20 / 2149.8Source ↗official
SafetyBenchIA#19 / 2162.4Source ↗official
SafetyBenchMH#18 / 2170.2Source ↗official
SafetyBenchOFF#18 / 2148.9Source ↗official
SafetyBenchPH#18 / 2154.5Source ↗official
SafetyBenchPP#18 / 2165Source ↗official
SafetyBenchUB#9 / 2163.2Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#13 / 3320.47Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#14 / 3317.59Source ↗official
SALAD-Benchattack_enhanced_malicious_use#14 / 3316.31Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#13 / 3318.75Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#14 / 3318.21Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#8 / 3323.81Source ↗official
SALAD-Benchbase_human_autonomy_integrity#5 / 3398.37Source ↗official
SALAD-Benchbase_information_safety_harms#3 / 3399.26Source ↗official
SALAD-Benchbase_malicious_use#4 / 3398.54Source ↗official
SALAD-Benchbase_misinformation_harms#9 / 3396.26Source ↗official
SALAD-Benchbase_representation_toxicity#9 / 3393.06Source ↗official
SALAD-Benchbase_socioeconomic_harms#6 / 3394.12Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#27 / 336.389Source ↗official
SALAD-Benchmcq_information_safety_harms#29 / 333.333Source ↗official
SALAD-Benchmcq_malicious_use#28 / 334.551Source ↗official
SALAD-Benchmcq_misinformation_harms#27 / 336.19Source ↗official
SALAD-Benchmcq_representation_toxicity#28 / 334.271Source ↗official
SALAD-Benchmcq_socioeconomic_harms#29 / 331.667Source ↗official
SORRY-Benchavg#12 / 510.14Source ↗official