← Models

Model profile

Qwen 2.5 14B Instruct

Alibabadeveloper
2024-09-19release date
#235 / 309overall rank
6eval lineages

Evidence summary

Qwen 2.5 14B Instruct has an estimated overall rank of #235; its 90% source-sensitivity interval is #136–#254. Its behavior-only rank is #227; company governance moves the combined estimate to #235. Published evidence spans 6 evals and 5 of 7 behavior components. Its strongest relative result is PandaBench JBB direct-request panel (safety_rate, #1 of 46); its weakest is SafeDialBench (legality, #18 of 18).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Agent-SafetyBenchcompromise_availability#10 / 1629.2Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#12 / 1629.2Source ↗official
Agent-SafetyBenchleak_sensitive_information#13 / 1624.4Source ↗official
Agent-SafetyBenchphysical_harm#11 / 1628Source ↗official
Agent-SafetyBenchproduce_unsafe_information#12 / 1681.2Source ↗official
Agent-SafetyBenchproperty_loss#11 / 1631.2Source ↗official
Agent-SafetyBenchspread_unsafe_information#10 / 1611.2Source ↗official
Agent-SafetyBenchviolate_law_ethics#12 / 1620.4Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#165 / 24112.92Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#105 / 24189.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#120 / 24169.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#85 / 23997.73Source ↗official
M3-SafetyBenchoverall_score#3 / 1995.94Source ↗official
PandaBench JBB direct-request panelsafety_rate#1 / 461Source ↗official
SafeDialBenchaggression#10 / 187.123Source ↗official
SafeDialBenchethics#14 / 187.39Source ↗official
SafeDialBenchfairness#5 / 187.56Source ↗official
SafeDialBenchlegality#18 / 187.17Source ↗official
SafeDialBenchmorality#17 / 187.047Source ↗official
SafeDialBenchprivacy#15 / 187.117Source ↗official
SYCON Benchfalse_presupposition_tof#7 / 112.31Source ↗official
SYCON Benchunethical_queries_tof#9 / 111.14Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-14.3
Government50.9
Diplomacy63.3
Economy45.3
Society60