← Models

Model profile

Qwen1.5 14B Chat

Alibabadeveloper
2024-02-04release date
#141 / 267overall rank
6eval lineages

Evidence summary

Qwen1.5 14B Chat has an estimated overall rank of #141; its 90% source-sensitivity interval is #57–#198. Its behavior-only rank is #134; company governance moves the combined estimate to #141. Published evidence spans 6 evals and 5 of 7 behavior components. Its strongest relative result is CRiskEval (malicious_coordination, #2 of 17); its weakest is SALAD-Bench (attack_enhanced_misinformation_harms, #31 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
ChineseSafescore#7 / 2261.29↑ higherSource ↗official
ChiSafetyBenchharmful_response_rate#2 / 140.22↓ lowerSource ↗official
ChiSafetyBenchmcq_score#3 / 1288.39↑ higherSource ↗official
CRiskEvaldeception_willingness#11 / 1723.53↓ lowerSource ↗official
CRiskEvaldesire_for_resource#5 / 1728.02↓ lowerSource ↗official
CRiskEvalharmful_goal#5 / 1734.36↓ lowerSource ↗official
CRiskEvalimprovement_intent#5 / 1746.04↓ lowerSource ↗official
CRiskEvalmalicious_coordination#2 / 176.52↓ lowerSource ↗official
CRiskEvalself_preservation#5 / 1729.41↓ lowerSource ↗official
CRiskEvalsituational_awareness#12 / 1766.62↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#97 / 26019.12↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#165 / 26086.5↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#166 / 26057.78↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#237 / 25880.95↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#26 / 337.11↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#30 / 333.91↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#26 / 336.53↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#31 / 334.28↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#26 / 336.19↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#22 / 339.09↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#14 / 3396.1↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#8 / 3397.77↑ higherSource ↗official
SALAD-Benchbase_malicious_use#10 / 3396.87↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#10 / 3395.72↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#11 / 3392.71↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#8 / 3393.65↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#6 / 3362.22↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#9 / 3349.72↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#7 / 3355.77↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#8 / 3354.52↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#7 / 3357.71↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#7 / 3352.78↑ higherSource ↗official
SORRY-Benchavg#30 / 510.34↓ lowerSource ↗official