Model profile
Evidence summary
Qwen 2.5 7B Instruct has an estimated overall rank of #227; its 90% source-sensitivity interval is #77–#272. Its behavior-only rank is #219; company governance moves the combined estimate to #227. Published evidence spans 20 evals and 7 of 7 behavior components. Its strongest relative result is TAC (base_welfare_rate, #6 of 76); its weakest is Agent-SafetyBench (harmful_vulnerable_code, #16 of 16).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Hugging Face ↗Qwen/Qwen2.5-7B-Instruct
- OpenRouter ↗qwen/qwen-2.5-7b-instruct
- Official model page ↗Exact model document · Reviewed official Hugging Face owner · official repository
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Source |
|---|---|---|---|---|
| Agent-SafetyBenchcompromise_availability | #15 / 16 | ↑17.2 | Source ↗official | |
| Agent-SafetyBenchharmful_vulnerable_code | #16 / 16 | ↑10.8 | Source ↗official | |
| Agent-SafetyBenchleak_sensitive_information | #15 / 16 | ↑13.2 | Source ↗official | |
| Agent-SafetyBenchphysical_harm | #15 / 16 | ↑17.6 | Source ↗official | |
| Agent-SafetyBenchproduce_unsafe_information | #15 / 16 | ↑57.6 | Source ↗official | |
| Agent-SafetyBenchproperty_loss | #15 / 16 | ↑15.6 | Source ↗official | |
| Agent-SafetyBenchspread_unsafe_information | #13 / 16 | ↑7.6 | Source ↗official | |
| Agent-SafetyBenchviolate_law_ethics | #15 / 16 | ↑10.4 | Source ↗official | |
| AgentDrive Safety Compliancescr | #40 / 48 | ↑53.75 | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #67 / 80 | ↑0.47 | Source ↗official | |
| DSPSafeBenchscore | #5 / 12 | ↑73.51 | Source ↗official | |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | #40 / 241 | ↑35.53 | Source ↗official | |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | #178 / 241 | ↑83.5 | Source ↗official | |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | #191 / 241 | ↑45.27 | Source ↗official | |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | #222 / 239 | ↑78.73 | Source ↗official | |
| FinEval Financial Security Knowledgefinancial_security_accuracy_pct | #12 / 19 | ↑71.7 | Source ↗official | |
| HELM Safetyanthropic_red_team | #49 / 80 | ↑0.985 | Source ↗official | |
| HELM Safetybbq | #54 / 80 | ↑0.906 | Source ↗official | |
| HELM Safetyharmbench | #48 / 80 | ↑0.677 | Source ↗official | |
| HELM Safetysimple_safety_tests | #65 / 80 | ↑0.96 | Source ↗official | |
| HELM Safetyxstest | #30 / 80 | ↑0.966 | Source ↗official | |
| IndoBias-Pairs — parity-aware culturally grounded biasparity_score | #11 / 26 | ↑86.3 | Source ↗official | |
| M3-SafetyBenchoverall_score | #8 / 19 | ↑92.37 | Source ↗official | |
| PandaBench JBB direct-request panelsafety_rate | #14 / 46 | ↑0.99 | Source ↗official | |
| SafeDialBenchaggression | #16 / 18 | ↑7.013 | Source ↗official | |
| SafeDialBenchethics | #15 / 18 | ↑7.357 | Source ↗official | |
| SafeDialBenchfairness | #7 / 18 | ↑7.553 | Source ↗official | |
| SafeDialBenchlegality | #15 / 18 | ↑7.21 | Source ↗official | |
| SafeDialBenchmorality | #16 / 18 | ↑7.06 | Source ↗official | |
| SafeDialBenchprivacy | #18 / 18 | ↑7.05 | Source ↗official | |
| Shelleducation_jsr | #12 / 14 | ↓0.804 | Source ↗official | |
| Shellfinance_jsr | #14 / 14 | ↓0.914 | Source ↗official | |
| Shellmanagement_jsr | #14 / 14 | ↓0.938 | Source ↗official | |
| SYCON Benchfalse_presupposition_tof | #8 / 11 | ↑1.93 | Source ↗official | |
| SYCON Benchunethical_queries_tof | #11 / 11 | ↑0.72 | Source ↗official | |
| TACbase_welfare_rate | #6 / 76 | ↑46.2 | Source ↗self run | |
| ThaiSafetyBenchsafety_score | #7 / 18 | ↑85.57 | Source ↗official | |
| UAVBench safety-critical decision recognitionethical_safety_critical_accuracy | #26 / 27 | ↑0.535 | Source ↗official |
Values evaluations
Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.