← Models

Model profile

Qwen1.5 0.5B Chat

Alibabadeveloper
2024-02-04release date
#227 / 267overall rank
3eval lineages

Evidence summary

Qwen1.5 0.5B Chat has an estimated overall rank of #227; its 90% source-sensitivity interval is #42–#257. Its behavior-only rank is #223; company governance moves the combined estimate to #227. Published evidence spans 3 evals and 4 of 7 behavior components. Its strongest relative result is Open LLM Safety Index (jailbreakbench_safety_rate, #1 of 21); its weakest is SALAD-Bench (mcq_representation_toxicity, #31 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Open LLM Safety Indexjailbreakbench_safety_rate#1 / 210.8↑ higherSource ↗official
Open LLM Safety Indexstrongreject_safety_rate#1 / 211↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#20 / 3311.42↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#25 / 335.21↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#21 / 338.4↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#25 / 336.91↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#25 / 336.79↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#18 / 3312.12↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#27 / 3383.23↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#28 / 3384.9↑ higherSource ↗official
SALAD-Benchbase_malicious_use#27 / 3380.92↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#25 / 3388.08↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#28 / 3375.3↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#25 / 3381.2↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#30 / 330↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#30 / 330↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#31 / 330↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#30 / 330↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#31 / 330↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#30 / 330↑ higherSource ↗official
SORRY-Benchavg#41 / 510.43↓ lowerSource ↗official