← Models

Model profile

Mistral Medium 3.1

Mistral AIdeveloper
2025-08-12release date
#269 / 309overall rank
7eval lineages

Evidence summary

Mistral Medium 3.1 has an estimated overall rank of #269; its 90% source-sensitivity interval is #164–#297. Its behavior-only rank is #259; company governance moves the combined estimate to #269. Published evidence spans 7 evals and 6 of 7 behavior components. Its strongest relative result is AgentDrive Safety Compliance (scr, #2 of 48); its weakest is SpeciEval (sea_animal_4ns, #111 of 113).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#188 / 3280.8458Source ↗official
AgentDrive Safety Compliancescr#2 / 4896.25Source ↗official
Alignment Leaderboardcorrigibility#17 / 244.117Source ↗official
Alignment Leaderboardhonesty#20 / 243.212Source ↗official
Alignment Leaderboardnon_manipulation#22 / 242.665Source ↗official
Alignment Leaderboardrobustness#19 / 243.147Source ↗official
Alignment Leaderboardsafety#23 / 242.962Source ↗official
Alignment Leaderboardscheming#21 / 243.319Source ↗official
PHAREbias_resistance_diagnostic#45 / 660.3966Source ↗official
PHAREhallucination_resistance_diagnostic#46 / 700.6963Source ↗official
PHAREharm_resistance_diagnostic#40 / 700.9232Source ↗official
PHAREjailbreak_resistance_diagnostic#63 / 670.3083Source ↗official
SpeciEvalbelief_animal_sentience#37 / 1136.88Source ↗official
SpeciEvalland_animal_4ns#100 / 1134.97Source ↗official
SpeciEvalsea_animal_4ns#111 / 1135.58Source ↗official
SpeciEvalspeciesism#64 / 1132.09Source ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#6 / 270.75Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#92 / 9477.3Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20
Government42.6
Diplomacy66.8
Economy43.8
Society61.1