← Models

Model profile

Minimax M2

MiniMaxdeveloper
2025-10-27release date
#174 / 309overall rank
6eval lineages

Evidence summary

Minimax M2 has an estimated overall rank of #174; its 90% source-sensitivity interval is #55–#275. Its behavior-only rank is #175; company governance moves the combined estimate to #174. Published evidence spans 6 evals and 6 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (bias_attack_non_success_rate, #14 of 241); its weakest is SpeciEval (sea_animal_4ns, #105 of 113).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#264 / 3280.909Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#96 / 1121405.0Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#32 / 10418.05Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#14 / 24151.94Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#156 / 24186Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#49 / 24189.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#36 / 23999.09Source ↗official
LiveSecBenchethics#12 / 4369.21Source ↗official
LiveSecBenchfactuality#13 / 4362.29Source ↗official
LiveSecBenchlegality#12 / 4367.96Source ↗official
LiveSecBenchprivacy#14 / 4364.5Source ↗official
LiveSecBenchpsychological_health#10 / 4369.47Source ↗official
SpeciEvalbelief_animal_sentience#54 / 1136.82Source ↗official
SpeciEvalland_animal_4ns#100 / 1134.97Source ↗official
SpeciEvalsea_animal_4ns#105 / 1135.2Source ↗official
SpeciEvalspeciesism#40 / 1131.81Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-16.9
Government45.2
Diplomacy64
Economy48.6
Society59.9

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility55
Warning-phrase rate35
Soft-censorship rate0
API-error rate0

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions4.01
CCP-narrative alignment — China topics4.52
CCP-narrative alignment — non-China controls2.47