← Models

Model profile

Mimo V2 Flash

Xiaomideveloper
2025-12-16release date
#157 / 267overall rank
4eval lineages

Evidence summary

Mimo V2 Flash has an estimated overall rank of #157; its 90% source-sensitivity interval is #45–#222. Its behavior-only rank is #158; company governance moves the combined estimate to #157. Published evidence spans 4 evals and 5 of 7 behavior components. Its strongest relative result is LiveSecBench (ethics, #6 of 43); its weakest is SM-Bench (overfit, #72 of 73).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#46 / 3110.4839↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#75 / 1050.145↑ higherSource ↗official
LiveSecBenchethics#6 / 4380.99↑ higherSource ↗official
LiveSecBenchfactuality#19 / 4352.77↑ higherSource ↗official
LiveSecBenchlegality#18 / 4352.92↑ higherSource ↗official
LiveSecBenchprivacy#22 / 4353.69↑ higherSource ↗official
LiveSecBenchpsychological_health#26 / 4345.77↑ higherSource ↗official
SM-Benchadversarial#64 / 7373.66↑ higherSource ↗official
SM-Benchambiguous_interpretation#50 / 7381.85↑ higherSource ↗official
SM-Benchanti_hallucination#62 / 7380.63↑ higherSource ↗official
SM-Bencheq_boundaries#40 / 7363.2↑ higherSource ↗official
SM-Benchoverfit#72 / 7312.84↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-29.9
Government45
Diplomacy66.2
Economy51.5
Society66.3