← Models

Model profile

Qwen 14B Chat

Alibabadeveloper
2023-09-25release date
#112 / 267overall rank
8eval lineages

Evidence summary

Qwen 14B Chat has an estimated overall rank of #112; its 90% source-sensitivity interval is #42–#186. Its behavior-only rank is #102; company governance moves the combined estimate to #112. Published evidence spans 8 evals and 4 of 7 behavior components. Its strongest relative result is CMoralEval (familial_morality, #2 of 26); its weakest is FLAMES (fairness, #11 of 13).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
CMoralEvalfamilial_morality#2 / 260.59↑ higherSource ↗official
CMoralEvalinternet_ethics#2 / 260.55↑ higherSource ↗official
CMoralEvalpersonal_morality#2 / 260.54↑ higherSource ↗official
CMoralEvalprofessional_ethics#2 / 260.57↑ higherSource ↗official
CMoralEvalsocial_morality#2 / 260.56↑ higherSource ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#3 / 1469.33↑ higherSource ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#3 / 1498.67↑ higherSource ↗official
FLAMESdata_protection#3 / 1355.26↑ higherSource ↗official
FLAMESfairness#11 / 1330.92↑ higherSource ↗official
FLAMESlegality#9 / 1332.61↑ higherSource ↗official
FLAMESmorality#1 / 1354.23↑ higherSource ↗official
FLAMESsafety#4 / 1336.83↑ higherSource ↗official
HarmBenchdr#10 / 2816.5↓ lowerSource ↗official
S-Evalbase_en_overall#6 / 2273.5↑ higherSource ↗official
SafeDialBenchaggression#7 / 187.15↑ higherSource ↗official
SafeDialBenchethics#2 / 187.68↑ higherSource ↗official
SafeDialBenchfairness#11 / 187.27↑ higherSource ↗official
SafeDialBenchlegality#4 / 187.987↑ higherSource ↗official
SafeDialBenchmorality#2 / 187.437↑ higherSource ↗official
SafeDialBenchprivacy#3 / 187.69↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#22 / 339.91↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#24 / 336.51↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#18 / 3310.44↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#23 / 338.39↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#22 / 337.44↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#24 / 337.79↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#8 / 3397.32↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#12 / 3396.34↑ higherSource ↗official
SALAD-Benchbase_malicious_use#8 / 3397.33↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#13 / 3395.42↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#13 / 3392.21↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#9 / 3393.07↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#7 / 3361.94↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#7 / 3351.11↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#8 / 3354.74↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#7 / 3355↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#8 / 3355.42↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#6 / 3355.56↑ higherSource ↗official
SuperCLUE Safetyinstruction_attack#9 / 3168.97↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#13 / 3161.82↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#21 / 3172.34↑ higherSource ↗official