← Models

Model profile

Llama 3 70B Instruct

Metadeveloper
2024-04-18release date
#232 / 333Safety rank
#183 / 645Freedom rank

Evidence summary

Safety. Llama 3 70B Instruct has an estimated Safety rank of #232; its 90% source-sensitivity interval is #123–#274. Its behavior-only rank is #216; company governance moves the combined estimate to #232. Published Safety evidence spans 13 eval lineages and 7 of 7 components. Its strongest relative result is Large-scale Moral Machine experiment on LLMs (human_choice_distance, #5 of 39); its weakest is AgentDojo (utility_under_attack, #15 of 15).

Freedom. Llama 3 70B Instruct has an estimated Freedom rank of #183; its 90% source-sensitivity interval is #158–#335. Published Freedom evidence spans 10 eval lineages and 1 of 1 components. Its strongest relative result is OR-Bench (toxic_acceptance_rate, #3 of 25); its weakest is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #205 of 246).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#242 / 3450.8754Source ↗official
AgentDojotargeted_attack_success_rate#11 / 150.256Source ↗official
AgentDojoutility_under_attack#15 / 150.1828Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#45 / 800.646Source ↗official
CASE-Benchagreement_accuracy#3 / 784.44Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#122 / 24816.02Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#126 / 24887.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#63 / 24887.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#42 / 24698.95Source ↗official
HELM Safetyanthropic_red_team#66 / 800.967Source ↗official
HELM Safetybbq#53 / 800.91Source ↗official
HELM Safetyharmbench#57 / 800.64Source ↗official
HELM Safetysimple_safety_tests#32 / 800.99Source ↗official
HELM Safetyxstest#29 / 800.968Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#5 / 390.7475Source ↗official
MT-JailBench CrescendoXsafety_score#17 / 2111.32Source ↗official
OR-Benchover_refusal_rate#9 / 2537.7Source ↗official
OR-Benchtoxic_acceptance_rate#22 / 2521.3Source ↗official
S-Evalbase_en_overall#15 / 2254.7Source ↗official
SORRY-Benchavg#31 / 510.35Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#36 / 800.646Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#121 / 24887.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#185 / 24887.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#205 / 24698.95Source ↗official
HELM Safetyanthropic_red_team#15 / 800.967Source ↗official
HELM Safetyharmbench#24 / 800.64Source ↗official
HELM Safetysimple_safety_tests#42 / 800.99Source ↗official
HELM Safetyxstest#29 / 800.968Source ↗official
MT-JailBench CrescendoXsafety_score#3 / 2111.32Source ↗official
OR-Benchover_refusal_rate#9 / 2537.7Source ↗official
OR-Benchtoxic_acceptance_rate#3 / 2521.3Source ↗official
S-Evalbase_en_overall#8 / 2254.7Source ↗official
SORRY-Benchavg#19 / 510.35Source ↗official
SpeechMap model completioncomplete_pct#55 / 18170.8Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism64.7
Self-direction47.2
Care / Harm12.3
Fairness / Cheating11.9
Ethical88.2