← Models

Model profile

Qwen 2.5 Max

Alibabadeveloper
2025-01-25release date
#259 / 309overall rank
4eval lineages

Evidence summary

Qwen 2.5 Max has an estimated overall rank of #259; its 90% source-sensitivity interval is #159–#293. Its behavior-only rank is #253; company governance moves the combined estimate to #259. Published evidence spans 4 evals and 6 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #45 of 241); its weakest is SM-Bench (eq_boundaries, #75 of 79).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Confabulationsconfabulation_rate#38 / 5231.19Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#112 / 24116.28Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#45 / 24192.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#127 / 24167.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#67 / 23998.09Source ↗official
PHAREbias_resistance_diagnostic#40 / 660.4295Source ↗official
PHAREhallucination_resistance_diagnostic#45 / 700.698Source ↗official
PHAREharm_resistance_diagnostic#50 / 700.8989Source ↗official
PHAREjailbreak_resistance_diagnostic#39 / 670.4626Source ↗official
SM-Benchadversarial#75 / 7970.73Source ↗official
SM-Benchambiguous_interpretation#65 / 7977.38Source ↗official
SM-Benchanti_hallucination#51 / 7988.48Source ↗official
SM-Bencheq_boundaries#75 / 7947.47Source ↗official
SM-Benchoverfit#46 / 7966.67Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism79.8
Self-direction56.1
Care / Harm28.6
Fairness / Cheating27.2
Ethical89.7