← Models

Model profile

Qwen1.5 72B Chat

Alibabadeveloper
2024-02-04release date
#174 / 312overall rank
13eval lineages

Evidence summary

Qwen1.5 72B Chat has an estimated overall rank of #174; its 90% source-sensitivity interval is #88–#215. Its behavior-only rank is #165; company governance moves the combined estimate to #174. Published evidence spans 13 evals and 6 of 7 behavior components. Its strongest relative result is SALAD-Bench (mcq_representation_toxicity, #2 of 33); its weakest is CRiskEval (situational_awareness, #15 of 17).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#66 / 800.486Source ↗official
ChineseSafescore#5 / 2263.67Source ↗official
ChiSafetyBenchharmful_response_rate#2 / 140.22Source ↗official
ChiSafetyBenchmcq_score#1 / 1291.13Source ↗official
COMPL-AI AI-Identity Disclosurescore#11 / 140.726Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#7 / 140.4856Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#8 / 130.4536Source ↗official
CRiskEvaldeception_willingness#6 / 1720.12Source ↗official
CRiskEvaldesire_for_resource#7 / 1731.71Source ↗official
CRiskEvalharmful_goal#4 / 1733.39Source ↗official
CRiskEvalimprovement_intent#7 / 1748.6Source ↗official
CRiskEvalmalicious_coordination#6 / 178.07Source ↗official
CRiskEvalself_preservation#8 / 1736.86Source ↗official
CRiskEvalsituational_awareness#15 / 1768.75Source ↗official
HELM Safetyanthropic_red_team#37 / 800.99Source ↗official
HELM Safetybbq#64 / 800.846Source ↗official
HELM Safetyharmbench#56 / 800.648Source ↗official
HELM Safetysimple_safety_tests#32 / 800.99Source ↗official
HELM Safetyxstest#42 / 800.957Source ↗official
OR-Benchover_refusal_rate#12 / 2546.9Source ↗official
OR-Benchtoxic_acceptance_rate#16 / 255.6Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#13 / 3320.47Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#13 / 3317.92Source ↗official
SALAD-Benchattack_enhanced_malicious_use#13 / 3317.05Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#14 / 3318.42Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#16 / 3314.19Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#15 / 3314.29Source ↗official
SALAD-Benchbase_human_autonomy_integrity#14 / 3396.1Source ↗official
SALAD-Benchbase_information_safety_harms#10 / 3396.89Source ↗official
SALAD-Benchbase_malicious_use#15 / 3395.2Source ↗official
SALAD-Benchbase_misinformation_harms#17 / 3393.65Source ↗official
SALAD-Benchbase_representation_toxicity#16 / 3390.06Source ↗official
SALAD-Benchbase_socioeconomic_harms#12 / 3391.89Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#2 / 3383.89Source ↗official
SALAD-Benchmcq_information_safety_harms#2 / 3380.28Source ↗official
SALAD-Benchmcq_malicious_use#2 / 3384.81Source ↗official
SALAD-Benchmcq_misinformation_harms#2 / 3380Source ↗official
SALAD-Benchmcq_representation_toxicity#2 / 3379.27Source ↗official
SALAD-Benchmcq_socioeconomic_harms#2 / 3378.89Source ↗official
SORRY-Benchavg#34 / 510.36Source ↗official