← Models

Model profile

MiMo v2 Pro

Xiaomideveloper
2026-03-18release date
#57 / 333Safety rank
#493 / 645Freedom rank

Evidence summary

Safety. MiMo v2 Pro has an estimated Safety rank of #57; its 90% source-sensitivity interval is #9–#192. Its behavior-only rank is #51; company governance moves the combined estimate to #57. Published Safety evidence spans 5 eval lineages and 5 of 7 components. Its strongest relative result is SM-Bench (anti_hallucination, #1 of 84); its weakest is Concordia — Shutdown-Resistance (safety_score, #38 of 53).

Freedom. MiMo v2 Pro has an estimated Freedom rank of #493; its 90% source-sensitivity interval is #111–#645. Published Freedom evidence spans 3 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (eq_boundaries, #4 of 84); its weakest is SpeechMap model completion (complete_pct, #178 of 181).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#31 / 3450.2998Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#57 / 1111435.0Source ↗official
Concordia — Agentic-Misalignmentsafety_score#37 / 5466.5Source ↗official
Concordia — AirBench-Deceptionsafety_score#18 / 6391.85Source ↗official
Concordia — AirBench-Manipulationsafety_score#19 / 5695.33Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#31 / 6393.69Source ↗official
Concordia — APEsafety_score#23 / 5535.62Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#9 / 6396.02Source ↗official
Concordia — DarkBenchsafety_score#6 / 5568.94Source ↗official
Concordia — Fortress-Biologicalsafety_score#24 / 5459.38Source ↗official
Concordia — Fortress-Chemicalsafety_score#21 / 5465.41Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#19 / 5470.07Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#13 / 4547.74Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#17 / 4536.5Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#20 / 453.333Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#8 / 4574.67Source ↗official
Concordia — MASKsafety_score#19 / 6276.73Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#40 / 6373.99Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#39 / 6316.11Source ↗official
Concordia — Shutdown-Resistancesafety_score#38 / 5395Source ↗official
Concordia — SOSBench-Biosafety_score#10 / 6396.6Source ↗official
Concordia — SOSBench-Chemsafety_score#17 / 6392.2Source ↗official
SM-Benchadversarial#9 / 8488.29Source ↗official
SM-Benchambiguous_interpretation#46 / 8485.12Source ↗official
SM-Benchanti_hallucination#1 / 84100Source ↗official
SM-Bencheq_boundaries#4 / 8479.78Source ↗official
SM-Benchoverfit#16 / 8490.71Source ↗official
WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct#15 / 2437.48Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Concordia — AirBench-Deceptionsafety_score#46 / 6391.85Source ↗official
Concordia — AirBench-Manipulationsafety_score#38 / 5695.33Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#35 / 5668.57Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#33 / 6393.69Source ↗official
Concordia — Fortress-Biologicalsafety_score#31 / 5459.38Source ↗official
Concordia — Fortress-Chemicalsafety_score#34 / 5465.41Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#36 / 5470.07Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#33 / 4547.74Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#29 / 4536.5Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#22 / 453.333Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#38 / 4574.67Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#24 / 6373.99Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#25 / 6316.11Source ↗official
Concordia — SOSBench-Biosafety_score#54 / 6396.6Source ↗official
Concordia — SOSBench-Chemsafety_score#47 / 6392.2Source ↗official
SM-Benchadversarial#74 / 8488.29Source ↗official
SM-Bencheq_boundaries#4 / 8479.78Source ↗official
SM-Benchoverfit#16 / 8490.71Source ↗official
SpeechMap model completioncomplete_pct#178 / 18125.4Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions3.13
CCP-narrative alignment — China topics3.62
CCP-narrative alignment — non-China controls1.66