← Models

Model profile

Qwen1.5 32B Chat

Alibabadeveloper
2024-04-02release date
#154 / 312overall rank
4eval lineages

Evidence summary

Qwen1.5 32B Chat has an estimated overall rank of #154; its 90% source-sensitivity interval is #84–#213. Its behavior-only rank is #149; company governance moves the combined estimate to #154. Published evidence spans 4 evals and 6 of 7 behavior components. Its strongest relative result is ChiSafetyBench (harmful_response_rate, #2 of 14); its weakest is CRiskEval (situational_awareness, #10 of 17).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChiSafetyBenchharmful_response_rate#2 / 140.22Source ↗official
ChiSafetyBenchmcq_score#4 / 1287.3Source ↗official
CRiskEvaldeception_willingness#5 / 1719.92Source ↗official
CRiskEvaldesire_for_resource#6 / 1728.56Source ↗official
CRiskEvalharmful_goal#7 / 1735.65Source ↗official
CRiskEvalimprovement_intent#8 / 1748.96Source ↗official
CRiskEvalmalicious_coordination#9 / 1710.48Source ↗official
CRiskEvalself_preservation#6 / 1729.81Source ↗official
CRiskEvalsituational_awareness#10 / 1765.38Source ↗official
OR-Benchover_refusal_rate#13 / 2550.8Source ↗official
OR-Benchtoxic_acceptance_rate#13 / 254.4Source ↗official
SORRY-Benchavg#25 / 510.28Source ↗official