← Models

Model profile

Qwen3.6 Plus

Alibabadeveloper
2026-04-02release date
#146 / 333Safety rank
#608 / 645Freedom rank

Evidence summary

Safety. Qwen3.6 Plus has an estimated Safety rank of #146; its 90% source-sensitivity interval is #86–#257. Its behavior-only rank is #140; company governance moves the combined estimate to #146. Published Safety evidence spans 11 eval lineages and 7 of 7 components. Its strongest relative result is BullshitBench v2 (clear_pushback_rate, #15 of 117); its weakest is Every Model Cheats — Cybench Cheat Propensity (standard_anti_cheat_cheat_propensity_pct, #21 of 22).

Freedom. Qwen3.6 Plus has an estimated Freedom rank of #608; its 90% source-sensitivity interval is #377–#640. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is UGI Leaderboard — base-model willingness (willingness_adherence_score, #70 of 156); its weakest is SpeechMap model completion (complete_pct, #157 of 181).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#45 / 3450.3464Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#9 / 3123.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#10 / 3155.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#12 / 3147.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#6 / 3126.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#11 / 3141.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#8 / 3140Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#9 / 3143.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#7 / 3133.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#9 / 3131.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#10 / 3146.73Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#10 / 3134Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#10 / 3152.2Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#85 / 1111418.0Source ↗official
BullshitBench v2clear_pushback_rate#15 / 1170.655Source ↗official
DystopiaBenchbasaglia_score#38 / 5069Source ↗official
DystopiaBenchbaudrillard_score#44 / 5073.63Source ↗official
DystopiaBenchhuxley_score#42 / 5077.27Source ↗official
DystopiaBenchlaguardia_score#33 / 5068.67Source ↗official
DystopiaBenchorwell_score#38 / 5073.9Source ↗official
DystopiaBenchpetrov_score#39 / 5078.1Source ↗official
Every Model Cheats — Cybench Cheat Propensitybaseline_cheat_propensity_pct#16 / 2243.5Source ↗official
Every Model Cheats — Cybench Cheat Propensitysevere_anti_cheat_cheat_propensity_pct#14 / 228.7Source ↗official
Every Model Cheats — Cybench Cheat Propensitystandard_anti_cheat_cheat_propensity_pct#21 / 2239.1Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#23 / 690.3Source ↗official
SM-Benchadversarial#21 / 8485.85Source ↗official
SM-Benchambiguous_interpretation#22 / 8489.29Source ↗official
SM-Benchanti_hallucination#26 / 8496.86Source ↗official
SM-Bencheq_boundaries#57 / 8458.71Source ↗official
SM-Benchoverfit#48 / 8469.4Source ↗official
SpeciEvalbelief_animal_sentience#21 / 1236.98Source ↗official
SpeciEvalland_animal_4ns#65 / 1234.55Source ↗official
SpeciEvalsea_animal_4ns#51 / 1234.67Source ↗official
SpeciEvalspeciesism#87 / 1232.27Source ↗official
The Dictatorship Evaloverall_resistance_rate#15 / 2038.83Source ↗official
ToolPrivacyBenchprivate_mt_poi#7 / 927.46Source ↗official
ToolPrivacyBenchpublic_mt_poi#7 / 919.25Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#23 / 3123.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#22 / 3155.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#17 / 3147.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#25 / 3126.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#21 / 3141.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#24 / 3140Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#23 / 3143.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#25 / 3133.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#23 / 3131.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#22 / 3146.73Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#22 / 3134Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#22 / 3152.2Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#45 / 690.3Source ↗official
SM-Benchadversarial#62 / 8485.85Source ↗official
SM-Bencheq_boundaries#57 / 8458.71Source ↗official
SM-Benchoverfit#48 / 8469.4Source ↗official
SpeechMap model completioncomplete_pct#157 / 18134.4Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#70 / 1562.25Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#99 / 1562.5Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20.2
Government47
Diplomacy65.4
Economy45.1
Society61.5