Model profile
Evidence summary
Safety. Mistral Large 2 has an estimated Safety rank of #318; its 90% source-sensitivity interval is #265–#324. Its behavior-only rank is #310; company governance moves the combined estimate to #318. Published Safety evidence spans 14 eval lineages and 5 of 7 components. Its strongest relative result is Vectara HHEM Factual Consistency (factual_consistency_rate, #6 of 94); its weakest is AgentHarm (harm_score, #12 of 12).
Freedom. Mistral Large 2 has an estimated Freedom rank of #6; its 90% source-sensitivity interval is #21–#87. Published Freedom evidence spans 11 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is LiveSecBench (legality, #20 of 43).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗mistral-large-2
- Hugging Face ↗mistralai/Mistral-Large-Instruct-2407
- OpenRouter ↗mistralai/mistral-large-2407
- Official model page ↗Exact model document · Reviewed official Hugging Face owner · official repository
- Release source ↗direct research preview identity date
Safety evals
Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Source |
|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | #120 / 345 | ↓0.6768 | Source ↗official | |
| Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr | #22 / 24 | ↓20.81 | Source ↗official | |
| Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr | #21 / 24 | ↓69.42 | Source ↗official | |
| AgentHarmharm_score | #12 / 12 | ↓82.2 | Source ↗official | |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | #28 / 32 | ↓35.9 | Source ↗official | |
| AILuminate General Purpose AI Chatdefamation | #29 / 32 | ↓39.2 | Source ↗official | |
| AILuminate General Purpose AI Chathate | #29 / 32 | ↓24.7 | Source ↗official | |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | #28 / 32 | ↓49.4 | Source ↗official | |
| AILuminate General Purpose AI Chatintellectual_property | #27 / 32 | ↓25.4 | Source ↗official | |
| AILuminate General Purpose AI Chatnon_violent_crimes | #29 / 32 | ↓39.4 | Source ↗official | |
| AILuminate General Purpose AI Chatprivacy | #29 / 32 | ↓31.7 | Source ↗official | |
| AILuminate General Purpose AI Chatsex_related_crimes | #28 / 32 | ↓34.4 | Source ↗official | |
| AILuminate General Purpose AI Chatsexual_content | #28 / 32 | ↓48.6 | Source ↗official | |
| AILuminate General Purpose AI Chatspecialized_advice | #29 / 32 | ↓33.8 | Source ↗official | |
| AILuminate General Purpose AI Chatsuicide_self_harm | #27 / 32 | ↓27.3 | Source ↗official | |
| AILuminate General Purpose AI Chatviolent_crimes | #28 / 32 | ↓45.7 | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #75 / 80 | ↑0.353 | Source ↗official | |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | #87 / 104 | ↓69.79 | Source ↗official | |
| Confabulationsconfabulation_rate | #39 / 52 | ↓32.18 | Source ↗official | |
| Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct | #27 / 69 | ↑0 | Source ↗official | |
| LiveSecBenchethics | #33 / 43 | ↑27.9 | Source ↗official | |
| LiveSecBenchfactuality | #37 / 43 | ↑26.5 | Source ↗official | |
| LiveSecBenchlegality | #24 / 43 | ↑42.83 | Source ↗official | |
| LiveSecBenchprivacy | #34 / 43 | ↑26.82 | Source ↗official | |
| LiveSecBenchpsychological_health | #37 / 43 | ↑24.54 | Source ↗official | |
| MASKlying_probability_pct | #46 / 53 | ↓52.47 | Source ↗official | |
| PHAREbias_resistance_diagnostic | #44 / 66 | ↑0.397 | Source ↗official | |
| PHAREhallucination_resistance_diagnostic | #33 / 70 | ↑0.7562 | Source ↗official | |
| PHAREharm_resistance_diagnostic | #52 / 70 | ↑0.8938 | Source ↗official | |
| PHAREjailbreak_resistance_diagnostic | #64 / 67 | ↑0.3058 | Source ↗official | |
| Shelleducation_jsr | #11 / 14 | ↓0.79 | Source ↗official | |
| Shellfinance_jsr | #13 / 14 | ↓0.912 | Source ↗official | |
| Shellmanagement_jsr | #13 / 14 | ↓0.92 | Source ↗official | |
| SORRY-Benchavg | #45 / 51 | ↓0.6 | Source ↗official | |
| Vectara HHEM Factual Consistencyfactual_consistency_rate | #6 / 94 | ↑95.5 | Source ↗official |
Freedom evals
Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Source |
|---|---|---|---|---|
| Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr | #3 / 24 | ↑20.81 | Source ↗official | |
| Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr | #4 / 24 | ↑69.42 | Source ↗official | |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | #5 / 32 | ↑35.9 | Source ↗official | |
| AILuminate General Purpose AI Chatdefamation | #4 / 32 | ↑39.2 | Source ↗official | |
| AILuminate General Purpose AI Chathate | #4 / 32 | ↑24.7 | Source ↗official | |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | #5 / 32 | ↑49.4 | Source ↗official | |
| AILuminate General Purpose AI Chatintellectual_property | #6 / 32 | ↑25.4 | Source ↗official | |
| AILuminate General Purpose AI Chatnon_violent_crimes | #4 / 32 | ↑39.4 | Source ↗official | |
| AILuminate General Purpose AI Chatprivacy | #4 / 32 | ↑31.7 | Source ↗official | |
| AILuminate General Purpose AI Chatsex_related_crimes | #5 / 32 | ↑34.4 | Source ↗official | |
| AILuminate General Purpose AI Chatsexual_content | #5 / 32 | ↑48.6 | Source ↗official | |
| AILuminate General Purpose AI Chatspecialized_advice | #4 / 32 | ↑33.8 | Source ↗official | |
| AILuminate General Purpose AI Chatsuicide_self_harm | #6 / 32 | ↑27.3 | Source ↗official | |
| AILuminate General Purpose AI Chatviolent_crimes | #5 / 32 | ↑45.7 | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #6 / 80 | ↓0.353 | Source ↗official | |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | #18 / 104 | ↑69.79 | Source ↗official | |
| Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct | #1 / 69 | ↓0 | Source ↗official | |
| LiveSecBenchethics | #11 / 43 | ↓27.9 | Source ↗official | |
| LiveSecBenchlegality | #20 / 43 | ↓42.83 | Source ↗official | |
| LiveSecBenchprivacy | #10 / 43 | ↓26.82 | Source ↗official | |
| LiveSecBenchpsychological_health | #7 / 43 | ↓24.54 | Source ↗official | |
| PHAREharm_resistance_diagnostic | #19 / 70 | ↓0.8938 | Source ↗official | |
| PHAREjailbreak_resistance_diagnostic | #4 / 67 | ↓0.3058 | Source ↗official | |
| Shelleducation_jsr | #4 / 14 | ↑0.79 | Source ↗official | |
| Shellfinance_jsr | #2 / 14 | ↑0.912 | Source ↗official | |
| Shellmanagement_jsr | #2 / 14 | ↑0.92 | Source ↗official | |
| SORRY-Benchavg | #7 / 51 | ↑0.6 | Source ↗official | |
| SpeechMap model completioncomplete_pct | #8 / 181 | ↑92 | Source ↗official | |
| UGI Leaderboard — base-model willingnesswillingness_adherence_score | #13 / 156 | ↑7.75 | Source ↗official | |
| UGI Leaderboard — base-model willingnesswillingness_direct_score | #21 / 156 | ↑5 | Source ↗official |
Values evaluations
Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.
