← Models

Model profile

MiniMax M2

MiniMaxdeveloper
2025-10-27release date
#131 / 333Safety rank
#499 / 645Freedom rank

Evidence summary

Safety. MiniMax M2 has an estimated Safety rank of #131; its 90% source-sensitivity interval is #62–#267. Its behavior-only rank is #134; company governance moves the combined estimate to #131. Published Safety evidence spans 7 eval lineages and 6 of 7 components. Its strongest relative result is Concordia — Shutdown-Resistance (safety_score, #1 of 53); its weakest is SpeciEval (sea_animal_4ns, #115 of 123).

Freedom. MiniMax M2 has an estimated Freedom rank of #499; its 90% source-sensitivity interval is #366–#545. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is Concordia — FRT-SciKnowEval-BiologicalHarmfulQA (safety_score, #11 of 45); its weakest is Concordia — Fortress-Privacy/Scams (safety_score, #48 of 54).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#281 / 3450.909Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#95 / 1111405.0Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#32 / 10418.05Source ↗official
Concordia — Agentic-Misalignmentsafety_score#23 / 5482.5Source ↗official
Concordia — AirBench-Deceptionsafety_score#19 / 6390.74Source ↗official
Concordia — AirBench-Manipulationsafety_score#12 / 5698Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#18 / 6398.2Source ↗official
Concordia — APEsafety_score#13 / 5558.36Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#51 / 6373.31Source ↗official
Concordia — DarkBenchsafety_score#10 / 5566.98Source ↗official
Concordia — Fortress-Biologicalsafety_score#12 / 5481.11Source ↗official
Concordia — Fortress-Chemicalsafety_score#14 / 5479.07Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#7 / 5479.81Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#11 / 4550Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#14 / 4538.67Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#33 / 451.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#12 / 4565Source ↗official
Concordia — MASKsafety_score#14 / 6280.51Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#32 / 6382.49Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#18 / 6332.82Source ↗official
Concordia — Shutdown-Resistancesafety_score#1 / 53100Source ↗official
Concordia — SOSBench-Biosafety_score#27 / 6387.6Source ↗official
Concordia — SOSBench-Chemsafety_score#33 / 6386.8Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#15 / 24851.94Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#155 / 24886Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#57 / 24889.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#36 / 24699.09Source ↗official
LiveSecBenchethics#12 / 4369.21Source ↗official
LiveSecBenchfactuality#13 / 4362.29Source ↗official
LiveSecBenchlegality#12 / 4367.96Source ↗official
LiveSecBenchprivacy#14 / 4364.5Source ↗official
LiveSecBenchpsychological_health#10 / 4369.47Source ↗official
SpeciEvalbelief_animal_sentience#59 / 1236.82Source ↗official
SpeciEvalland_animal_4ns#110 / 1234.97Source ↗official
SpeciEvalsea_animal_4ns#115 / 1235.2Source ↗official
SpeciEvalspeciesism#44 / 1231.81Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#73 / 10418.05Source ↗official
Concordia — AirBench-Deceptionsafety_score#45 / 6390.74Source ↗official
Concordia — AirBench-Manipulationsafety_score#41 / 5698Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#46 / 5681.9Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#46 / 6398.2Source ↗official
Concordia — Fortress-Biologicalsafety_score#43 / 5481.11Source ↗official
Concordia — Fortress-Chemicalsafety_score#41 / 5479.07Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#48 / 5479.81Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#35 / 4550Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#32 / 4538.67Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#11 / 451.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#34 / 4565Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#31 / 6382.49Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#45 / 6332.82Source ↗official
Concordia — SOSBench-Biosafety_score#37 / 6387.6Source ↗official
Concordia — SOSBench-Chemsafety_score#31 / 6386.8Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#92 / 24886Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#192 / 24889.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#210 / 24699.09Source ↗official
LiveSecBenchethics#32 / 4369.21Source ↗official
LiveSecBenchlegality#32 / 4367.96Source ↗official
LiveSecBenchprivacy#30 / 4364.5Source ↗official
LiveSecBenchpsychological_health#34 / 4369.47Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#61 / 1563Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#113 / 1562Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-16.9
Government45.2
Diplomacy64
Economy48.6
Society59.9

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility55
Warning-phrase rate35
Soft-censorship rate0
API-error rate0

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions4.01
CCP-narrative alignment — China topics4.52
CCP-narrative alignment — non-China controls2.47