← Models

Model profile

Minimax M2.5

MiniMaxdeveloper
2026-02-12release date
#139 / 312overall rank
13eval lineages

Evidence summary

Minimax M2.5 has an estimated overall rank of #139; its 90% source-sensitivity interval is #72–#236. Its behavior-only rank is #140; company governance moves the combined estimate to #139. Published evidence spans 13 evals and 7 of 7 behavior components. Its strongest relative result is StereoTales Harmful Associations (benign_significant_association_score, #1 of 23); its weakest is WildClawBench Safety & Alignment (OpenClaw harness) (safety_alignment_score_pct, #23 of 24).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#235 / 3300.8812Source ↗official
AgentAbstainabstain#14 / 1750.1Source ↗official
AgentAbstaincar#14 / 1749.6Source ↗official
AgentAbstainpaired#11 / 1741.9Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#49 / 1121440.0Source ↗official
BullshitBench v2clear_pushback_rate#90 / 1060.085Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#24 / 1049.521Source ↗official
DystopiaBenchbasaglia_score#11 / 5044.73Source ↗official
DystopiaBenchbaudrillard_score#8 / 5031.17Source ↗official
DystopiaBenchhuxley_score#18 / 5054.33Source ↗official
DystopiaBenchlaguardia_score#22 / 5063.47Source ↗official
DystopiaBenchorwell_score#14 / 5045.1Source ↗official
DystopiaBenchpetrov_score#19 / 5070.3Source ↗official
LiveSecBenchethics#10 / 4372.04Source ↗official
LiveSecBenchfactuality#5 / 4376.82Source ↗official
LiveSecBenchlegality#21 / 4348.6Source ↗official
LiveSecBenchprivacy#19 / 4359.81Source ↗official
LiveSecBenchpsychological_health#22 / 4350.99Source ↗official
SABERoverall_safety_rate#6 / 1326.34Source ↗official
SABERscenario_a_safety_rate#4 / 1332.83Source ↗official
SABERscenario_b_safety_rate#5 / 1334.83Source ↗official
SABERscenario_c_safety_rate#11 / 1312.23Source ↗official
SM-Benchadversarial#73 / 7973.17Source ↗official
SM-Benchambiguous_interpretation#53 / 7982.14Source ↗official
SM-Benchanti_hallucination#71 / 7977.49Source ↗official
SM-Bencheq_boundaries#49 / 7961.24Source ↗official
SM-Benchoverfit#61 / 7950.82Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#1 / 2389.25Source ↗official
TACbase_welfare_rate#59 / 7620.51Source ↗self run
Vectara HHEM Factual Consistencyfactual_consistency_rate#40 / 9490.9Source ↗official
WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct#23 / 2429.8Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.4
Government49.9
Diplomacy65.4
Economy42.7
Society56.9

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions3.78
CCP-narrative alignment — China topics4.29
CCP-narrative alignment — non-China controls2.25

Moral Trolley Arena