← Models

Model profile

Qwen3 14B

Alibabadeveloper
2025-04-27release date
#251 / 309overall rank
5eval lineages

Evidence summary

Qwen3 14B has an estimated overall rank of #251; its 90% source-sensitivity interval is #163–#275. Its behavior-only rank is #247; company governance moves the combined estimate to #251. Published evidence spans 5 evals and 6 of 7 behavior components. Its strongest relative result is Vectara HHEM Factual Consistency (factual_consistency_rate, #15 of 94); its weakest is Alignment Leaderboard (robustness, #23 of 24).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Alignment Leaderboardcorrigibility#13 / 244.204Source ↗official
Alignment Leaderboardhonesty#22 / 243.147Source ↗official
Alignment Leaderboardnon_manipulation#21 / 243.013Source ↗official
Alignment Leaderboardrobustness#23 / 242.813Source ↗official
Alignment Leaderboardsafety#19 / 243.558Source ↗official
Alignment Leaderboardscheming#18 / 243.458Source ↗official
MuPPET Contextual Privacymultiparty_contextual_privacy_score#5 / 735.47Source ↗official
PandaBench JBB direct-request panelsafety_rate#28 / 460.97Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#15 / 9494.6Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#18 / 2510.2Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20
Government49.1
Diplomacy64.9
Economy46.9
Society61.5