← Models

Model profile

Qwen3.6 Plus

Alibabadeveloper
2026-04-02release date
#131 / 309overall rank
7eval lineages

Evidence summary

Qwen3.6 Plus has an estimated overall rank of #131; its 90% source-sensitivity interval is #41–#238. Its behavior-only rank is #121; company governance moves the combined estimate to #131. Published evidence spans 7 evals and 6 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #31 of 328); its weakest is DystopiaBench (baudrillard_score, #44 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#31 / 3280.3464Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#86 / 1121418.0Source ↗official
BullshitBench v2clear_pushback_rate#12 / 1060.655Source ↗official
DystopiaBenchbasaglia_score#38 / 5069Source ↗official
DystopiaBenchbaudrillard_score#44 / 5073.63Source ↗official
DystopiaBenchhuxley_score#42 / 5077.27Source ↗official
DystopiaBenchlaguardia_score#33 / 5068.67Source ↗official
DystopiaBenchorwell_score#38 / 5073.9Source ↗official
DystopiaBenchpetrov_score#39 / 5078.1Source ↗official
SM-Benchadversarial#19 / 7985.85Source ↗official
SM-Benchambiguous_interpretation#19 / 7989.29Source ↗official
SM-Benchanti_hallucination#23 / 7996.86Source ↗official
SM-Bencheq_boundaries#52 / 7958.71Source ↗official
SM-Benchoverfit#43 / 7969.4Source ↗official
SpeciEvalbelief_animal_sentience#19 / 1136.98Source ↗official
SpeciEvalland_animal_4ns#58 / 1134.55Source ↗official
SpeciEvalsea_animal_4ns#45 / 1134.67Source ↗official
SpeciEvalspeciesism#79 / 1132.27Source ↗official
ToolPrivacyBenchprivate_mt_poi#7 / 927.46Source ↗official
ToolPrivacyBenchpublic_mt_poi#7 / 919.25Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20.2
Government47
Diplomacy65.4
Economy45.1
Society61.5