← Models

Model profile

Qwen 2.5 7B Instruct

Alibabadeveloper
2024-09-19release date
#198 / 267overall rank
16eval lineages

Evidence summary

Qwen 2.5 7B Instruct has an estimated overall rank of #198; its 90% source-sensitivity interval is #75–#247. Its behavior-only rank is #189; company governance moves the combined estimate to #198. Published evidence spans 16 evals and 7 of 7 behavior components. Its strongest relative result is TAC (base_welfare_rate, #6 of 68); its weakest is Agent-SafetyBench (harmful_vulnerable_code, #16 of 16).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Agent-SafetyBenchcompromise_availability#15 / 1617.2↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#16 / 1610.8↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#15 / 1613.2↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#15 / 1617.6↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#15 / 1657.6↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#15 / 1615.6↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#13 / 167.6↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#15 / 1610.4↑ higherSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#67 / 800.47↑ higherSource ↗official
DSPSafeBenchscore#5 / 1273.51↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#81 / 26021.71↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#201 / 26082.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#215 / 26042.22↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#252 / 25861.82↑ higherSource ↗official
HELM Safetyanthropic_red_team#49 / 800.985↑ higherSource ↗official
HELM Safetybbq#54 / 800.906↑ higherSource ↗official
HELM Safetyharmbench#48 / 800.677↑ higherSource ↗official
HELM Safetysimple_safety_tests#65 / 800.96↑ higherSource ↗official
HELM Safetyxstest#30 / 800.966↑ higherSource ↗official
M3-SafetyBenchoverall_score#8 / 1992.37↑ higherSource ↗official
PandaBench JBB direct-request panelsafety_rate#14 / 460.99↑ higherSource ↗official
SafeDialBenchaggression#16 / 187.013↑ higherSource ↗official
SafeDialBenchethics#15 / 187.357↑ higherSource ↗official
SafeDialBenchfairness#7 / 187.553↑ higherSource ↗official
SafeDialBenchlegality#15 / 187.21↑ higherSource ↗official
SafeDialBenchmorality#16 / 187.06↑ higherSource ↗official
SafeDialBenchprivacy#18 / 187.05↑ higherSource ↗official
Shelleducation_jsr#12 / 140.804↓ lowerSource ↗official
Shellfinance_jsr#14 / 140.914↓ lowerSource ↗official
Shellmanagement_jsr#14 / 140.938↓ lowerSource ↗official
SYCON Benchfalse_presupposition_tof#8 / 111.93↑ higherSource ↗official
SYCON Benchunethical_queries_tof#11 / 110.72↑ higherSource ↗official
TACbase_welfare_rate#6 / 6846.2↑ higherSource ↗self run
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#26 / 270.535↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-14.7
Government48.8
Diplomacy56
Economy48.8
Society59.2