Model profile
Evidence summary
Mistral Large 2 has an estimated overall rank of #293; its 90% source-sensitivity interval is #241–#301. Its behavior-only rank is #285; company governance moves the combined estimate to #293. Published evidence spans 12 evals and 5 of 7 behavior components. Its strongest relative result is Vectara HHEM Factual Consistency (factual_consistency_rate, #6 of 94); its weakest is AgentHarm (harm_score, #12 of 12).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗mistral-large-2
- Hugging Face ↗mistralai/Mistral-Large-Instruct-2407
- OpenRouter ↗mistralai/mistral-large-2407
- Official model page ↗Exact model document · Reviewed official Hugging Face owner · official repository
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Source |
|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | #101 / 328 | ↓0.6768 | Source ↗official | |
| AgentHarmharm_score | #12 / 12 | ↓82.2 | Source ↗official | |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | #28 / 32 | ↓35.9 | Source ↗official | |
| AILuminate General Purpose AI Chatdefamation | #29 / 32 | ↓39.2 | Source ↗official | |
| AILuminate General Purpose AI Chathate | #29 / 32 | ↓24.7 | Source ↗official | |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | #28 / 32 | ↓49.4 | Source ↗official | |
| AILuminate General Purpose AI Chatintellectual_property | #27 / 32 | ↓25.4 | Source ↗official | |
| AILuminate General Purpose AI Chatnon_violent_crimes | #29 / 32 | ↓39.4 | Source ↗official | |
| AILuminate General Purpose AI Chatprivacy | #29 / 32 | ↓31.7 | Source ↗official | |
| AILuminate General Purpose AI Chatsex_related_crimes | #28 / 32 | ↓34.4 | Source ↗official | |
| AILuminate General Purpose AI Chatsexual_content | #28 / 32 | ↓48.6 | Source ↗official | |
| AILuminate General Purpose AI Chatspecialized_advice | #29 / 32 | ↓33.8 | Source ↗official | |
| AILuminate General Purpose AI Chatsuicide_self_harm | #27 / 32 | ↓27.3 | Source ↗official | |
| AILuminate General Purpose AI Chatviolent_crimes | #28 / 32 | ↓45.7 | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #75 / 80 | ↑0.353 | Source ↗official | |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | #87 / 104 | ↓69.79 | Source ↗official | |
| Confabulationsconfabulation_rate | #39 / 52 | ↓32.18 | Source ↗official | |
| LiveSecBenchethics | #33 / 43 | ↑27.9 | Source ↗official | |
| LiveSecBenchfactuality | #37 / 43 | ↑26.5 | Source ↗official | |
| LiveSecBenchlegality | #24 / 43 | ↑42.83 | Source ↗official | |
| LiveSecBenchprivacy | #34 / 43 | ↑26.82 | Source ↗official | |
| LiveSecBenchpsychological_health | #37 / 43 | ↑24.54 | Source ↗official | |
| MASKlying_probability_pct | #46 / 53 | ↓52.47 | Source ↗official | |
| PHAREbias_resistance_diagnostic | #44 / 66 | ↑0.397 | Source ↗official | |
| PHAREhallucination_resistance_diagnostic | #33 / 70 | ↑0.7562 | Source ↗official | |
| PHAREharm_resistance_diagnostic | #52 / 70 | ↑0.8938 | Source ↗official | |
| PHAREjailbreak_resistance_diagnostic | #64 / 67 | ↑0.3058 | Source ↗official | |
| Shelleducation_jsr | #11 / 14 | ↓0.79 | Source ↗official | |
| Shellfinance_jsr | #13 / 14 | ↓0.912 | Source ↗official | |
| Shellmanagement_jsr | #13 / 14 | ↓0.92 | Source ↗official | |
| SORRY-Benchavg | #45 / 51 | ↓0.6 | Source ↗official | |
| Vectara HHEM Factual Consistencyfactual_consistency_rate | #6 / 94 | ↑95.5 | Source ↗official |
Values evaluations
Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.
