← Models

Model profile

Minimax M2.5

MiniMaxdeveloper
2026-02-12release date
#144 / 267overall rank
9eval lineages

Evidence summary

Minimax M2.5 has an estimated overall rank of #144; its 90% source-sensitivity interval is #69–#218. Published evidence spans 9 evals and 7 of 7 behavior components. Its strongest relative result is LiveSecBench (factuality, #5 of 43); its weakest is SM-Bench (adversarial, #67 of 73).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#247 / 3110.893↓ lowerSource ↗official
AgentAbstainabstain#14 / 1750.1↑ higherSource ↗official
AgentAbstaincar#14 / 1749.6↑ higherSource ↗official
AgentAbstainpaired#11 / 1741.9↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#89 / 1050.085↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#24 / 1059.521↓ lowerSource ↗official
DystopiaBenchbasaglia_score#11 / 5044.73↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#8 / 5031.17↓ lowerSource ↗official
DystopiaBenchhuxley_score#18 / 5054.33↓ lowerSource ↗official
DystopiaBenchlaguardia_score#22 / 5063.47↓ lowerSource ↗official
DystopiaBenchorwell_score#14 / 5045.1↓ lowerSource ↗official
DystopiaBenchpetrov_score#19 / 5070.3↓ lowerSource ↗official
LiveSecBenchethics#10 / 4372.04↑ higherSource ↗official
LiveSecBenchfactuality#5 / 4376.82↑ higherSource ↗official
LiveSecBenchlegality#21 / 4348.6↑ higherSource ↗official
LiveSecBenchprivacy#19 / 4359.81↑ higherSource ↗official
LiveSecBenchpsychological_health#22 / 4350.99↑ higherSource ↗official
SABERoverall_safety_rate#6 / 1326.34↑ higherSource ↗official
SABERscenario_a_safety_rate#4 / 1332.83↑ higherSource ↗official
SABERscenario_b_safety_rate#5 / 1334.83↑ higherSource ↗official
SABERscenario_c_safety_rate#11 / 1312.23↑ higherSource ↗official
SM-Benchadversarial#67 / 7373.17↑ higherSource ↗official
SM-Benchambiguous_interpretation#48 / 7382.14↑ higherSource ↗official
SM-Benchanti_hallucination#66 / 7377.49↑ higherSource ↗official
SM-Bencheq_boundaries#44 / 7361.24↑ higherSource ↗official
SM-Benchoverfit#56 / 7350.82↑ higherSource ↗official
TACbase_welfare_rate#55 / 6820.51↑ higherSource ↗self run

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.4
Government49.9
Diplomacy65.4
Economy42.7
Society56.9