← Models

Model profile

Qwen 2.5 14B Instruct

Alibabadeveloper
2024-09-19release date
#201 / 267overall rank
6eval lineages

Evidence summary

Qwen 2.5 14B Instruct has an estimated overall rank of #201; its 90% source-sensitivity interval is #103–#227. Its behavior-only rank is #193; company governance moves the combined estimate to #201. Published evidence spans 6 evals and 5 of 7 behavior components. Its strongest relative result is PandaBench JBB direct-request panel (safety_rate, #1 of 46); its weakest is SafeDialBench (legality, #18 of 18).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Agent-SafetyBenchcompromise_availability#10 / 1629.2↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#12 / 1629.2↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#13 / 1624.4↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#11 / 1628↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#12 / 1681.2↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#11 / 1631.2↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#10 / 1611.2↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#12 / 1620.4↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#181 / 26012.92↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#116 / 26089.17↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#135 / 26069.44↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#96 / 25897.73↑ higherSource ↗official
M3-SafetyBenchoverall_score#3 / 1995.94↑ higherSource ↗official
PandaBench JBB direct-request panelsafety_rate#1 / 461↑ higherSource ↗official
SafeDialBenchaggression#10 / 187.123↑ higherSource ↗official
SafeDialBenchethics#14 / 187.39↑ higherSource ↗official
SafeDialBenchfairness#5 / 187.56↑ higherSource ↗official
SafeDialBenchlegality#18 / 187.17↑ higherSource ↗official
SafeDialBenchmorality#17 / 187.047↑ higherSource ↗official
SafeDialBenchprivacy#15 / 187.117↑ higherSource ↗official
SYCON Benchfalse_presupposition_tof#7 / 112.31↑ higherSource ↗official
SYCON Benchunethical_queries_tof#9 / 111.14↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-14.3
Government50.9
Diplomacy63.3
Economy45.3
Society60