← Models

Model profile

Mistral Large

Mistral AIdeveloper
2024-02-26release date
#77 / 267overall rank
6eval lineages

Evidence summary

Mistral Large has an estimated overall rank of #77; its 90% source-sensitivity interval is #28–#224. Its behavior-only rank is #64; company governance moves the combined estimate to #77. Published evidence spans 6 evals and 6 of 7 behavior components. Its strongest relative result is AnimalHarmBench (score, #1 of 10); its weakest is OR-Bench (toxic_acceptance_rate, #25 of 25).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Adversarial Robustnessscore#7 / 837↓ lowerSource ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#17 / 3221.5↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#24 / 3226.3↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#18 / 329.1↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#22 / 3231.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#21 / 3218↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#26 / 3231.1↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#25 / 3225.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#23 / 3221.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#22 / 3234.3↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#23 / 3222.7↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#14 / 3212↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#22 / 3228.7↓ lowerSource ↗official
AnimalHarmBenchscore#1 / 100.068↑ higherSource ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#2 / 1888.04↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#80 / 10563.06↓ lowerSource ↗official
OR-Benchover_refusal_rate#3 / 259.8↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#25 / 2527.2↓ lowerSource ↗official