← Models

Model profile

Magistral Medium

Mistral AIdeveloper
2025-06-10release date
#225 / 333Safety rank
#5 / 645Freedom rank

Evidence summary

Safety. Magistral Medium has an estimated Safety rank of #225; its 90% source-sensitivity interval is #121–#315. Its behavior-only rank is #212; company governance moves the combined estimate to #225. Published Safety evidence spans 8 eval lineages and 6 of 7 components. Its strongest relative result is FlagEval Safety and Values (a3_qualified_rate, #3 of 18); its weakest is Adversarial Poetry — AILuminate Baseline and Poetry ASR (baseline_asr, #24 of 24).

Freedom. Magistral Medium has an estimated Freedom rank of #5; its 90% source-sensitivity interval is #15–#103. Published Freedom evidence spans 4 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is PHARE (jailbreak_resistance_diagnostic, #27 of 67).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#93 / 3450.5966Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#24 / 2422.92Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#24 / 2477.19Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#127 / 24815.76Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#180 / 24883.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#230 / 24833.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#194 / 24691.82Source ↗official
FlagEval Safety and Valuesa1_qualified_rate#12 / 1879.01Source ↗official
FlagEval Safety and Valuesa2_qualified_rate#12 / 1878.73Source ↗official
FlagEval Safety and Valuesa3_qualified_rate#3 / 1889.28Source ↗official
FlagEval Safety and Valuesa4_qualified_rate#7 / 1889.15Source ↗official
FlagEval Safety and Valuesa5_qualified_rate#3 / 1874.27Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#36 / 5426.12Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 690Source ↗official
MASKlying_probability_pct#31 / 5343.5Source ↗official
PHAREbias_resistance_diagnostic#25 / 660.5075Source ↗official
PHAREhallucination_resistance_diagnostic#49 / 700.6819Source ↗official
PHAREharm_resistance_diagnostic#60 / 700.8452Source ↗official
PHAREjailbreak_resistance_diagnostic#41 / 670.4561Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#1 / 2422.92Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#1 / 2477.19Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#69 / 24883.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#18 / 24833.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#52 / 24691.82Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 690Source ↗official
PHAREharm_resistance_diagnostic#11 / 700.8452Source ↗official
PHAREjailbreak_resistance_diagnostic#27 / 670.4561Source ↗official