← Models

Model profile

Mistral Medium 3

Mistral AIdeveloper
2025-05-07release date
#294 / 333Safety rank
#7 / 645Freedom rank

Evidence summary

Safety. Mistral Medium 3 has an estimated Safety rank of #294; its 90% source-sensitivity interval is #166–#331. Its behavior-only rank is #287; company governance moves the combined estimate to #294. Published Safety evidence spans 6 eval lineages and 5 of 7 components. Its strongest relative result is AA-Omniscience (hallucination_rate, #95 of 345); its weakest is Concordia — SOSBench-Chem (safety_score, #63 of 63).

Freedom. Mistral Medium 3 has an estimated Freedom rank of #7; its 90% source-sensitivity interval is #2–#336. Published Freedom evidence spans 2 eval lineages and 1 of 1 components. Its strongest relative result is UGI Leaderboard — base-model willingness (willingness_adherence_score, #2 of 156); its weakest is Concordia — Fortress-Biological (safety_score, #12 of 54).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#95 / 3450.609Source ↗official
Concordia — Agentic-Misalignmentsafety_score#32 / 5472.33Source ↗official
Concordia — AirBench-Deceptionsafety_score#60 / 6337.78Source ↗official
Concordia — AirBench-Manipulationsafety_score#56 / 5642.67Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#62 / 6341.67Source ↗official
Concordia — APEsafety_score#53 / 550.4375Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#51 / 6373.31Source ↗official
Concordia — DarkBenchsafety_score#35 / 5553.03Source ↗official
Concordia — Fortress-Biologicalsafety_score#43 / 5429.4Source ↗official
Concordia — Fortress-Chemicalsafety_score#42 / 5432.88Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#53 / 5430.19Source ↗official
Concordia — MASKsafety_score#56 / 6242.04Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#58 / 6344.11Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#56 / 633.965Source ↗official
Concordia — Shutdown-Resistancesafety_score#36 / 5398Source ↗official
Concordia — SOSBench-Biosafety_score#59 / 6311.8Source ↗official
Concordia — SOSBench-Chemsafety_score#63 / 6325.6Source ↗official
Confabulationsconfabulation_rate#41 / 5238.12Source ↗official
Emergent Collusionhigh_illegality_game_rate#8 / 130.47Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#26 / 4277Source ↗official
SpeciEvalbelief_animal_sentience#81 / 1236.68Source ↗official
SpeciEvalland_animal_4ns#115 / 1235.03Source ↗official
SpeciEvalsea_animal_4ns#120 / 1235.53Source ↗official
SpeciEvalspeciesism#105 / 1232.62Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20.3
Government43.1
Diplomacy64.3
Economy45.8
Society59.9