← Models

Model profile

Qwen3.6 Max Preview

Alibabadeveloper
2026-04-22release date
#90 / 309overall rank
4eval lineages

Evidence summary

Qwen3.6 Max Preview has an estimated overall rank of #90; its 90% source-sensitivity interval is #43–#254. Its behavior-only rank is #84; company governance moves the combined estimate to #90. Published evidence spans 4 evals and 3 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #50 of 328); its weakest is DystopiaBench (laguardia_score, #39 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#50 / 3280.4619Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#72 / 1121428.0Source ↗official
DystopiaBenchbasaglia_score#36 / 5068.63Source ↗official
DystopiaBenchbaudrillard_score#38 / 5066.57Source ↗official
DystopiaBenchhuxley_score#38 / 5075.8Source ↗official
DystopiaBenchlaguardia_score#39 / 5069.4Source ↗official
DystopiaBenchorwell_score#34 / 5073.17Source ↗official
DystopiaBenchpetrov_score#27 / 5073.17Source ↗official
ODCV-Benchaverage_severity#4 / 121.175Source ↗official
ODCV-Benchmisalignment_rate#5 / 1228.75Source ↗official