← Models

Model profile

Qwen 2.5 7B Instruct

Alibabadeveloper
2024-09-19release date
#227 / 309overall rank
20eval lineages

Evidence summary

Qwen 2.5 7B Instruct has an estimated overall rank of #227; its 90% source-sensitivity interval is #77–#272. Its behavior-only rank is #219; company governance moves the combined estimate to #227. Published evidence spans 20 evals and 7 of 7 behavior components. Its strongest relative result is TAC (base_welfare_rate, #6 of 76); its weakest is Agent-SafetyBench (harmful_vulnerable_code, #16 of 16).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Agent-SafetyBenchcompromise_availability#15 / 1617.2Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#16 / 1610.8Source ↗official
Agent-SafetyBenchleak_sensitive_information#15 / 1613.2Source ↗official
Agent-SafetyBenchphysical_harm#15 / 1617.6Source ↗official
Agent-SafetyBenchproduce_unsafe_information#15 / 1657.6Source ↗official
Agent-SafetyBenchproperty_loss#15 / 1615.6Source ↗official
Agent-SafetyBenchspread_unsafe_information#13 / 167.6Source ↗official
Agent-SafetyBenchviolate_law_ethics#15 / 1610.4Source ↗official
AgentDrive Safety Compliancescr#40 / 4853.75Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#67 / 800.47Source ↗official
DSPSafeBenchscore#5 / 1273.51Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#40 / 24135.53Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#178 / 24183.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#191 / 24145.27Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#222 / 23978.73Source ↗official
FinEval Financial Security Knowledgefinancial_security_accuracy_pct#12 / 1971.7Source ↗official
HELM Safetyanthropic_red_team#49 / 800.985Source ↗official
HELM Safetybbq#54 / 800.906Source ↗official
HELM Safetyharmbench#48 / 800.677Source ↗official
HELM Safetysimple_safety_tests#65 / 800.96Source ↗official
HELM Safetyxstest#30 / 800.966Source ↗official
IndoBias-Pairs — parity-aware culturally grounded biasparity_score#11 / 2686.3Source ↗official
M3-SafetyBenchoverall_score#8 / 1992.37Source ↗official
PandaBench JBB direct-request panelsafety_rate#14 / 460.99Source ↗official
SafeDialBenchaggression#16 / 187.013Source ↗official
SafeDialBenchethics#15 / 187.357Source ↗official
SafeDialBenchfairness#7 / 187.553Source ↗official
SafeDialBenchlegality#15 / 187.21Source ↗official
SafeDialBenchmorality#16 / 187.06Source ↗official
SafeDialBenchprivacy#18 / 187.05Source ↗official
Shelleducation_jsr#12 / 140.804Source ↗official
Shellfinance_jsr#14 / 140.914Source ↗official
Shellmanagement_jsr#14 / 140.938Source ↗official
SYCON Benchfalse_presupposition_tof#8 / 111.93Source ↗official
SYCON Benchunethical_queries_tof#11 / 110.72Source ↗official
TACbase_welfare_rate#6 / 7646.2Source ↗self run
ThaiSafetyBenchsafety_score#7 / 1885.57Source ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#26 / 270.535Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-14.7
Government48.8
Diplomacy56
Economy48.8
Society59.2