← Models

Model profile

Llama 3 70B Instruct

Metadeveloper
2024-04-18release date
#212 / 309overall rank
13eval lineages

Evidence summary

Llama 3 70B Instruct has an estimated overall rank of #212; its 90% source-sensitivity interval is #91–#248. Its behavior-only rank is #197; company governance moves the combined estimate to #212. Published evidence spans 13 evals and 7 of 7 behavior components. Its strongest relative result is Large-scale Moral Machine experiment on LLMs (human_choice_distance, #5 of 39); its weakest is AgentDojo (utility_under_attack, #15 of 15).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#224 / 3280.8754Source ↗official
AgentDojotargeted_attack_success_rate#11 / 150.256Source ↗official
AgentDojoutility_under_attack#15 / 150.1828Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#45 / 800.646Source ↗official
CASE-Benchagreement_accuracy#3 / 784.44Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#116 / 24116.02Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#127 / 24187.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#55 / 24187.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#42 / 23998.95Source ↗official
HELM Safetyanthropic_red_team#66 / 800.967Source ↗official
HELM Safetybbq#53 / 800.91Source ↗official
HELM Safetyharmbench#57 / 800.64Source ↗official
HELM Safetysimple_safety_tests#32 / 800.99Source ↗official
HELM Safetyxstest#29 / 800.968Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#5 / 390.7475Source ↗official
MT-JailBench CrescendoXsafety_score#17 / 2111.32Source ↗official
OR-Benchover_refusal_rate#9 / 2537.7Source ↗official
OR-Benchtoxic_acceptance_rate#22 / 2521.3Source ↗official
S-Evalbase_en_overall#15 / 2254.7Source ↗official
SORRY-Benchavg#31 / 510.35Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism64.7
Self-direction47.2
Care / Harm12.3
Fairness / Cheating11.9
Ethical88.2