Model profile
Evidence summary
Llama 3.1 8B Instruct has an estimated overall rank of #268; its 90% source-sensitivity interval is #195–#274. Its behavior-only rank is #257; company governance moves the combined estimate to #268. Published evidence spans 31 evals and 7 of 7 behavior components. Its strongest relative result is AgentHarm (harm_score, #1 of 12); its weakest is Agent-SafetyBench (leak_sensitive_information, #16 of 16).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗llama-3-1-instruct-8b
- NVIDIA ↗qc69jvmznzxy/llama-3_1-8b-instruct
- OpenRouter ↗meta-llama/llama-3.1-8b-instruct
- Official model page ↗Family-level model document · meta · first party
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Source |
|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | #45 / 328 | ↓0.4304 | Source ↗official | |
| AbstentionBenchanswer_unknown_f1 | #12 / 20 | ↑0.8667 | Source ↗official | |
| AbstentionBenchfalse_premise_f1 | #14 / 20 | ↑0.6532 | Source ↗official | |
| AbstentionBenchstale_f1 | #13 / 20 | ↑0.6218 | Source ↗official | |
| AbstentionBenchsubjective_f1 | #9 / 20 | ↑0.7449 | Source ↗official | |
| AbstentionBenchunderspecified_context_f1 | #9 / 20 | ↑0.656 | Source ↗official | |
| AbstentionBenchunderspecified_intent_f1 | #10 / 20 | ↑0.7338 | Source ↗official | |
| Agent-SafetyBenchcompromise_availability | #16 / 16 | ↑12.8 | Source ↗official | |
| Agent-SafetyBenchharmful_vulnerable_code | #14 / 16 | ↑24.8 | Source ↗official | |
| Agent-SafetyBenchleak_sensitive_information | #16 / 16 | ↑10 | Source ↗official | |
| Agent-SafetyBenchphysical_harm | #16 / 16 | ↑11.2 | Source ↗official | |
| Agent-SafetyBenchproduce_unsafe_information | #14 / 16 | ↑74.8 | Source ↗official | |
| Agent-SafetyBenchproperty_loss | #16 / 16 | ↑12.4 | Source ↗official | |
| Agent-SafetyBenchspread_unsafe_information | #15 / 16 | ↑6.4 | Source ↗official | |
| Agent-SafetyBenchviolate_law_ethics | #16 / 16 | ↑6.8 | Source ↗official | |
| AgentDrive Safety Compliancescr | #42 / 48 | ↑50 | Source ↗official | |
| AgentHarmharm_score | #1 / 12 | ↓3.1 | Source ↗official | |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | #15 / 32 | ↓18.9 | Source ↗official | |
| AILuminate General Purpose AI Chatdefamation | #15 / 32 | ↓15.8 | Source ↗official | |
| AILuminate General Purpose AI Chathate | #14 / 32 | ↓7.5 | Source ↗official | |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | #16 / 32 | ↓21.9 | Source ↗official | |
| AILuminate General Purpose AI Chatintellectual_property | #18 / 32 | ↓14.2 | Source ↗official | |
| AILuminate General Purpose AI Chatnon_violent_crimes | #16 / 32 | ↓16 | Source ↗official | |
| AILuminate General Purpose AI Chatprivacy | #16 / 32 | ↓12.2 | Source ↗official | |
| AILuminate General Purpose AI Chatsex_related_crimes | #13 / 32 | ↓11.2 | Source ↗official | |
| AILuminate General Purpose AI Chatsexual_content | #18 / 32 | ↓24.1 | Source ↗official | |
| AILuminate General Purpose AI Chatspecialized_advice | #13 / 32 | ↓14.3 | Source ↗official | |
| AILuminate General Purpose AI Chatsuicide_self_harm | #12 / 32 | ↓10.6 | Source ↗official | |
| AILuminate General Purpose AI Chatviolent_crimes | #15 / 32 | ↓17.2 | Source ↗official | |
| AIMS Safety-Classifier Competenceaverage_harmful_f1 | #9 / 11 | ↑0.749 | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #47 / 80 | ↑0.623 | Source ↗official | |
| BullshitBench v2clear_pushback_rate | #78 / 106 | ↑0.14 | Source ↗official | |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | #65 / 104 | ↓52.74 | Source ↗official | |
| Contextual MoralChoicehuman_agreement | #16 / 22 | ↑0.35 | Source ↗official | |
| DSPSafeBenchscore | #12 / 12 | ↑61.51 | Source ↗official | |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | #96 / 241 | ↑18.35 | Source ↗official | |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | #174 / 241 | ↑83.91 | Source ↗official | |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | #108 / 241 | ↑73.33 | Source ↗official | |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | #183 / 239 | ↑92.07 | Source ↗official | |
| HELM Safetyanthropic_red_team | #65 / 80 | ↑0.968 | Source ↗official | |
| HELM Safetybbq | #68 / 80 | ↑0.785 | Source ↗official | |
| HELM Safetyharmbench | #61 / 80 | ↑0.616 | Source ↗official | |
| HELM Safetysimple_safety_tests | #40 / 80 | ↑0.988 | Source ↗official | |
| HELM Safetyxstest | #48 / 80 | ↑0.953 | Source ↗official | |
| IndoBias-Pairs — parity-aware culturally grounded biasparity_score | #16 / 26 | ↑85.01 | Source ↗official | |
| KIDBench Implicit Child Cueimplicit_child_cue_total_mean | #12 / 13 | ↑3.11 | Source ↗official | |
| Large-scale Moral Machine experiment on LLMshuman_choice_distance | #37 / 39 | ↓1.623 | Source ↗official | |
| MuPPET Contextual Privacymultiparty_contextual_privacy_score | #6 / 7 | ↑35.12 | Source ↗official | |
| PandaBench JBB direct-request panelsafety_rate | #25 / 46 | ↑0.98 | Source ↗official | |
| PHAREbias_resistance_diagnostic | #38 / 66 | ↑0.4415 | Source ↗official | |
| PHAREhallucination_resistance_diagnostic | #60 / 70 | ↑0.6381 | Source ↗official | |
| PHAREharm_resistance_diagnostic | #62 / 70 | ↑0.8306 | Source ↗official | |
| PHAREjailbreak_resistance_diagnostic | #26 / 67 | ↑0.5881 | Source ↗official | |
| PropensityBenchscore | #10 / 14 | ↓66.5 | Source ↗official | |
| Shelleducation_jsr | #8 / 14 | ↓0.658 | Source ↗official | |
| Shellfinance_jsr | #9 / 14 | ↓0.6 | Source ↗official | |
| Shellmanagement_jsr | #10 / 14 | ↓0.724 | Source ↗official | |
| SORRY-Benchavg | #12 / 51 | ↓0.14 | Source ↗official | |
| SYCON Benchfalse_presupposition_tof | #11 / 11 | ↑1.45 | Source ↗official | |
| SYCON Benchunethical_queries_tof | #10 / 11 | ↑0.85 | Source ↗official | |
| ThaiSafetyBenchsafety_score | #17 / 18 | ↑71.76 | Source ↗official | |
| UAVBench safety-critical decision recognitionethical_safety_critical_accuracy | #24 / 27 | ↑0.57 | Source ↗official | |
| VETO Misfired Alignmentmisfired_alignment_rate_pct | #7 / 25 | ↓6.2 | Source ↗official |
Values evaluations
Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.