← Models

Model profile

Baichuan 2 13B Chat

Baichuandeveloper
2023-09-06release date
#148 / 309overall rank
13eval lineages

Evidence summary

Baichuan 2 13B Chat has an estimated overall rank of #148; its 90% source-sensitivity interval is #71–#210. Its behavior-only rank is #154; company governance moves the combined estimate to #148. Published evidence spans 13 evals and 5 of 7 behavior components. Its strongest relative result is SuperCLUE Safety (traditional_safety, #1 of 31); its weakest is SafetyBench (UB, #20 of 21).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChineseSafescore#1 / 2270.43Source ↗official
ChiSafetyBenchharmful_response_rate#5 / 140.43Source ↗official
ChiSafetyBenchmcq_score#7 / 1279.83Source ↗official
CMoralEvalfamilial_morality#13 / 260.39Source ↗official
CMoralEvalinternet_ethics#19 / 260.34Source ↗official
CMoralEvalpersonal_morality#17 / 260.36Source ↗official
CMoralEvalprofessional_ethics#19 / 260.36Source ↗official
CMoralEvalsocial_morality#16 / 260.37Source ↗official
CRiskEvaldeception_willingness#15 / 1737.75Source ↗official
CRiskEvaldesire_for_resource#11 / 1735.35Source ↗official
CRiskEvalharmful_goal#15 / 1754.94Source ↗official
CRiskEvalimprovement_intent#10 / 1750.82Source ↗official
CRiskEvalmalicious_coordination#12 / 1718.46Source ↗official
CRiskEvalself_preservation#15 / 1743.45Source ↗official
CRiskEvalsituational_awareness#7 / 1763.34Source ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#9 / 1445.33Source ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#1 / 14100Source ↗official
FinEval Financial Security Knowledgefinancial_security_accuracy_pct#16 / 1961.6Source ↗official
FLAMESdata_protection#9 / 1339.47Source ↗official
FLAMESfairness#6 / 1338.55Source ↗official
FLAMESlegality#7 / 1339.13Source ↗official
FLAMESmorality#6 / 1344.78Source ↗official
FLAMESsafety#2 / 1353.85Source ↗official
HarmBenchdr#14 / 2819.3Source ↗official
M3-SafetyBenchoverall_score#13 / 1986.39Source ↗official
S-Evalbase_en_overall#4 / 2277.4Source ↗official
SafeDialBenchaggression#12 / 187.03Source ↗official
SafeDialBenchethics#7 / 187.6Source ↗official
SafeDialBenchfairness#15 / 187.17Source ↗official
SafeDialBenchlegality#4 / 187.987Source ↗official
SafeDialBenchmorality#6 / 187.303Source ↗official
SafeDialBenchprivacy#5 / 187.617Source ↗official
SafetyBenchEM#5 / 2175.75Source ↗official
SafetyBenchIA#4 / 2182.65Source ↗official
SafetyBenchMH#7 / 2183.65Source ↗official
SafetyBenchOFF#6 / 2169.25Source ↗official
SafetyBenchPH#5 / 2176.35Source ↗official
SafetyBenchPP#4 / 2182.05Source ↗official
SafetyBenchUB#20 / 2149.2Source ↗official
SuperCLUE Safetyinstruction_attack#3 / 3172.41Source ↗official
SuperCLUE Safetyresponsible_ai#4 / 3172.73Source ↗official
SuperCLUE Safetytraditional_safety#1 / 3187.23Source ↗official