← Models

Model profile

MiMo v2 Flash

Xiaomideveloper
2025-12-16release date
#224 / 333Safety rank
#256 / 645Freedom rank

Evidence summary

Safety. MiMo v2 Flash has an estimated Safety rank of #224; its 90% source-sensitivity interval is #120–#281. Its behavior-only rank is #227; company governance moves the combined estimate to #224. Published Safety evidence spans 7 eval lineages and 6 of 7 components. Its strongest relative result is Concordia — Shutdown-Resistance (safety_score, #1 of 53); its weakest is SM-Bench (overfit, #83 of 84).

Freedom. MiMo v2 Flash has an estimated Freedom rank of #256; its 90% source-sensitivity interval is #117–#431. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (adversarial, #9 of 84); its weakest is SM-Bench (overfit, #83 of 84).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#71 / 3450.4841Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#60 / 1111433.0Source ↗official
BullshitBench v2clear_pushback_rate#87 / 1170.145Source ↗official
Concordia — Agentic-Misalignmentsafety_score#18 / 5487.58Source ↗official
Concordia — AirBench-Deceptionsafety_score#45 / 6372.22Source ↗official
Concordia — AirBench-Manipulationsafety_score#37 / 5680.67Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#53 / 6373.87Source ↗official
Concordia — APEsafety_score#46 / 553.347Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#56 / 6370.12Source ↗official
Concordia — DarkBenchsafety_score#38 / 5549.85Source ↗official
Concordia — Fortress-Biologicalsafety_score#49 / 5426.82Source ↗official
Concordia — Fortress-Chemicalsafety_score#40 / 5433.67Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#47 / 5438.52Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#31 / 4526.89Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#31 / 4517.17Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#33 / 451.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#32 / 4539.33Source ↗official
Concordia — MASKsafety_score#32 / 6259.47Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#57 / 6347.3Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#42 / 6312.78Source ↗official
Concordia — Shutdown-Resistancesafety_score#1 / 53100Source ↗official
Concordia — SOSBench-Biosafety_score#47 / 6359.92Source ↗official
Concordia — SOSBench-Chemsafety_score#54 / 6360.4Source ↗official
LiveSecBenchethics#6 / 4380.99Source ↗official
LiveSecBenchfactuality#19 / 4352.77Source ↗official
LiveSecBenchlegality#18 / 4352.92Source ↗official
LiveSecBenchprivacy#22 / 4353.69Source ↗official
LiveSecBenchpsychological_health#26 / 4345.77Source ↗official
SM-Benchadversarial#75 / 8473.66Source ↗official
SM-Benchambiguous_interpretation#60 / 8481.85Source ↗official
SM-Benchanti_hallucination#72 / 8480.63Source ↗official
SM-Bencheq_boundaries#48 / 8463.2Source ↗official
SM-Benchoverfit#83 / 8412.84Source ↗official
WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct#19 / 2433.55Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Concordia — AirBench-Deceptionsafety_score#19 / 6372.22Source ↗official
Concordia — AirBench-Manipulationsafety_score#20 / 5680.67Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#13 / 5644.76Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#11 / 6373.87Source ↗official
Concordia — Fortress-Biologicalsafety_score#6 / 5426.82Source ↗official
Concordia — Fortress-Chemicalsafety_score#15 / 5433.67Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#8 / 5438.52Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#15 / 4526.89Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#15 / 4517.17Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#11 / 451.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#14 / 4539.33Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#7 / 6347.3Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#22 / 6312.78Source ↗official
Concordia — SOSBench-Biosafety_score#17 / 6359.92Source ↗official
Concordia — SOSBench-Chemsafety_score#10 / 6360.4Source ↗official
LiveSecBenchethics#38 / 4380.99Source ↗official
LiveSecBenchlegality#26 / 4352.92Source ↗official
LiveSecBenchprivacy#22 / 4353.69Source ↗official
LiveSecBenchpsychological_health#18 / 4345.77Source ↗official
SM-Benchadversarial#9 / 8473.66Source ↗official
SM-Bencheq_boundaries#48 / 8463.2Source ↗official
SM-Benchoverfit#83 / 8412.84Source ↗official
SpeechMap model completioncomplete_pct#108 / 18150.4Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#54 / 1563.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#46 / 1564Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-29.9
Government45
Diplomacy66.2
Economy51.5
Society66.3

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions4.45
CCP-narrative alignment — China topics4.76
CCP-narrative alignment — non-China controls3.49