← Models

Model profile

Baichuan 2 13B Chat

Baichuandeveloper
2023-09-06release date
#135 / 267overall rank
12eval lineages

Evidence summary

Baichuan 2 13B Chat has an estimated overall rank of #135; its 90% source-sensitivity interval is #65–#184. Its behavior-only rank is #137; company governance moves the combined estimate to #135. Published evidence spans 12 evals and 5 of 7 behavior components. Its strongest relative result is SuperCLUE Safety (traditional_safety, #1 of 31); its weakest is SafetyBench (UB, #20 of 21).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
ChineseSafescore#1 / 2270.43↑ higherSource ↗official
ChiSafetyBenchharmful_response_rate#5 / 140.43↓ lowerSource ↗official
ChiSafetyBenchmcq_score#7 / 1279.83↑ higherSource ↗official
CMoralEvalfamilial_morality#13 / 260.39↑ higherSource ↗official
CMoralEvalinternet_ethics#19 / 260.34↑ higherSource ↗official
CMoralEvalpersonal_morality#17 / 260.36↑ higherSource ↗official
CMoralEvalprofessional_ethics#19 / 260.36↑ higherSource ↗official
CMoralEvalsocial_morality#16 / 260.37↑ higherSource ↗official
CRiskEvaldeception_willingness#15 / 1737.75↓ lowerSource ↗official
CRiskEvaldesire_for_resource#11 / 1735.35↓ lowerSource ↗official
CRiskEvalharmful_goal#15 / 1754.94↓ lowerSource ↗official
CRiskEvalimprovement_intent#10 / 1750.82↓ lowerSource ↗official
CRiskEvalmalicious_coordination#12 / 1718.46↓ lowerSource ↗official
CRiskEvalself_preservation#15 / 1743.45↓ lowerSource ↗official
CRiskEvalsituational_awareness#7 / 1763.34↓ lowerSource ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#9 / 1445.33↑ higherSource ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#1 / 14100↑ higherSource ↗official
FLAMESdata_protection#9 / 1339.47↑ higherSource ↗official
FLAMESfairness#6 / 1338.55↑ higherSource ↗official
FLAMESlegality#7 / 1339.13↑ higherSource ↗official
FLAMESmorality#6 / 1344.78↑ higherSource ↗official
FLAMESsafety#2 / 1353.85↑ higherSource ↗official
HarmBenchdr#14 / 2819.3↓ lowerSource ↗official
M3-SafetyBenchoverall_score#13 / 1986.39↑ higherSource ↗official
S-Evalbase_en_overall#4 / 2277.4↑ higherSource ↗official
SafeDialBenchaggression#12 / 187.03↑ higherSource ↗official
SafeDialBenchethics#7 / 187.6↑ higherSource ↗official
SafeDialBenchfairness#15 / 187.17↑ higherSource ↗official
SafeDialBenchlegality#4 / 187.987↑ higherSource ↗official
SafeDialBenchmorality#6 / 187.303↑ higherSource ↗official
SafeDialBenchprivacy#5 / 187.617↑ higherSource ↗official
SafetyBenchEM#5 / 2175.75↑ higherSource ↗official
SafetyBenchIA#4 / 2182.65↑ higherSource ↗official
SafetyBenchMH#7 / 2183.65↑ higherSource ↗official
SafetyBenchOFF#6 / 2169.25↑ higherSource ↗official
SafetyBenchPH#5 / 2176.35↑ higherSource ↗official
SafetyBenchPP#4 / 2182.05↑ higherSource ↗official
SafetyBenchUB#20 / 2149.2↑ higherSource ↗official
SuperCLUE Safetyinstruction_attack#3 / 3172.41↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#4 / 3172.73↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#1 / 3187.23↑ higherSource ↗official