← Models

Model profile

Qwen1.5 32B Chat

Alibabadeveloper
2024-04-02release date
#176 / 333Safety rank
#413 / 645Freedom rank

Evidence summary

Safety. Qwen1.5 32B Chat has an estimated Safety rank of #176; its 90% source-sensitivity interval is #99–#224. Its behavior-only rank is #168; company governance moves the combined estimate to #176. Published Safety evidence spans 4 eval lineages and 6 of 7 components. Its strongest relative result is ChiSafetyBench (harmful_response_rate, #2 of 14); its weakest is CRiskEval (situational_awareness, #10 of 17).

Freedom. Qwen1.5 32B Chat has an estimated Freedom rank of #413; its 90% source-sensitivity interval is #245–#589. Published Freedom evidence spans 3 eval lineages and 1 of 1 components. Its strongest relative result is OR-Bench (over_refusal_rate, #13 of 25); its weakest is ChiSafetyBench (refusal_rr1, #14 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChiSafetyBenchharmful_response_rate#2 / 140.22Source ↗official
ChiSafetyBenchmcq_score#4 / 1287.3Source ↗official
CRiskEvaldeception_willingness#5 / 1719.92Source ↗official
CRiskEvaldesire_for_resource#6 / 1728.56Source ↗official
CRiskEvalharmful_goal#7 / 1735.65Source ↗official
CRiskEvalimprovement_intent#8 / 1748.96Source ↗official
CRiskEvalmalicious_coordination#9 / 1710.48Source ↗official
CRiskEvalself_preservation#6 / 1729.81Source ↗official
CRiskEvalsituational_awareness#10 / 1765.38Source ↗official
OR-Benchover_refusal_rate#13 / 2550.8Source ↗official
OR-Benchtoxic_acceptance_rate#13 / 254.4Source ↗official
SORRY-Benchavg#25 / 510.28Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChiSafetyBenchharmful_response_rate#11 / 140.22Source ↗official
ChiSafetyBenchrefusal_rr1#14 / 1477.71Source ↗official
ChiSafetyBenchrefusal_rr2#14 / 1477.27Source ↗official
OR-Benchover_refusal_rate#13 / 2550.8Source ↗official
OR-Benchtoxic_acceptance_rate#13 / 254.4Source ↗official
SORRY-Benchavg#27 / 510.28Source ↗official