← Models

Model profile

Baichuan 2 7B Chat

Baichuandeveloper
2023-09-06release date
#220 / 309overall rank
11eval lineages

Evidence summary

Baichuan 2 7B Chat has an estimated overall rank of #220; its 90% source-sensitivity interval is #155–#246. Its behavior-only rank is #224; company governance moves the combined estimate to #220. Published evidence spans 11 evals and 5 of 7 behavior components. Its strongest relative result is FLAMES (safety, #1 of 13); its weakest is CRiskEval (improvement_intent, #17 of 17).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChineseSafescore#13 / 2253.99Source ↗official
ChiSafetyBenchharmful_response_rate#8 / 140.65Source ↗official
ChiSafetyBenchmcq_score#11 / 1250.03Source ↗official
CMoralEvalfamilial_morality#8 / 260.47Source ↗official
CMoralEvalinternet_ethics#6 / 260.48Source ↗official
CMoralEvalpersonal_morality#6 / 260.47Source ↗official
CMoralEvalprofessional_ethics#6 / 260.49Source ↗official
CMoralEvalsocial_morality#6 / 260.5Source ↗official
CRiskEvaldeception_willingness#16 / 1738.4Source ↗official
CRiskEvaldesire_for_resource#14 / 1739.06Source ↗official
CRiskEvalharmful_goal#14 / 1752.04Source ↗official
CRiskEvalimprovement_intent#17 / 1760.08Source ↗official
CRiskEvalmalicious_coordination#16 / 1729.29Source ↗official
CRiskEvalself_preservation#13 / 1740.2Source ↗official
CRiskEvalsituational_awareness#6 / 1761.78Source ↗official
DSPSafeBenchscore#11 / 1265.31Source ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#12 / 1420Source ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#5 / 1497.33Source ↗official
FLAMESdata_protection#8 / 1340.79Source ↗official
FLAMESfairness#4 / 1342.17Source ↗official
FLAMESlegality#4 / 1352.17Source ↗official
FLAMESmorality#11 / 1339.3Source ↗official
FLAMESsafety#1 / 1356.41Source ↗official
HarmBenchdr#13 / 2818.8Source ↗official
M3-SafetyBenchoverall_score#16 / 1982.91Source ↗official
SafeDialBenchaggression#11 / 187.073Source ↗official
SafeDialBenchethics#6 / 187.613Source ↗official
SafeDialBenchfairness#16 / 187.123Source ↗official
SafeDialBenchlegality#9 / 187.937Source ↗official
SafeDialBenchmorality#4 / 187.383Source ↗official
SafeDialBenchprivacy#11 / 187.523Source ↗official
SuperCLUE Safetyinstruction_attack#15 / 3165.52Source ↗official
SuperCLUE Safetyresponsible_ai#13 / 3161.82Source ↗official
SuperCLUE Safetytraditional_safety#5 / 3180.85Source ↗official