← Models

Model profile

Baichuan 2 7B Chat

Baichuandeveloper
2023-09-06release date
#191 / 267overall rank
11eval lineages

Evidence summary

Baichuan 2 7B Chat has an estimated overall rank of #191; its 90% source-sensitivity interval is #119–#217. Its behavior-only rank is #194; company governance moves the combined estimate to #191. Published evidence spans 11 evals and 5 of 7 behavior components. Its strongest relative result is FLAMES (safety, #1 of 13); its weakest is CRiskEval (improvement_intent, #17 of 17).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
ChineseSafescore#13 / 2253.99↑ higherSource ↗official
ChiSafetyBenchharmful_response_rate#8 / 140.65↓ lowerSource ↗official
ChiSafetyBenchmcq_score#11 / 1250.03↑ higherSource ↗official
CMoralEvalfamilial_morality#8 / 260.47↑ higherSource ↗official
CMoralEvalinternet_ethics#6 / 260.48↑ higherSource ↗official
CMoralEvalpersonal_morality#6 / 260.47↑ higherSource ↗official
CMoralEvalprofessional_ethics#6 / 260.49↑ higherSource ↗official
CMoralEvalsocial_morality#6 / 260.5↑ higherSource ↗official
CRiskEvaldeception_willingness#16 / 1738.4↓ lowerSource ↗official
CRiskEvaldesire_for_resource#14 / 1739.06↓ lowerSource ↗official
CRiskEvalharmful_goal#14 / 1752.04↓ lowerSource ↗official
CRiskEvalimprovement_intent#17 / 1760.08↓ lowerSource ↗official
CRiskEvalmalicious_coordination#16 / 1729.29↓ lowerSource ↗official
CRiskEvalself_preservation#13 / 1740.2↓ lowerSource ↗official
CRiskEvalsituational_awareness#6 / 1761.78↓ lowerSource ↗official
DSPSafeBenchscore#11 / 1265.31↑ higherSource ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#12 / 1420↑ higherSource ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#5 / 1497.33↑ higherSource ↗official
FLAMESdata_protection#8 / 1340.79↑ higherSource ↗official
FLAMESfairness#4 / 1342.17↑ higherSource ↗official
FLAMESlegality#4 / 1352.17↑ higherSource ↗official
FLAMESmorality#11 / 1339.3↑ higherSource ↗official
FLAMESsafety#1 / 1356.41↑ higherSource ↗official
HarmBenchdr#13 / 2818.8↓ lowerSource ↗official
M3-SafetyBenchoverall_score#16 / 1982.91↑ higherSource ↗official
SafeDialBenchaggression#11 / 187.073↑ higherSource ↗official
SafeDialBenchethics#6 / 187.613↑ higherSource ↗official
SafeDialBenchfairness#16 / 187.123↑ higherSource ↗official
SafeDialBenchlegality#9 / 187.937↑ higherSource ↗official
SafeDialBenchmorality#4 / 187.383↑ higherSource ↗official
SafeDialBenchprivacy#11 / 187.523↑ higherSource ↗official
SuperCLUE Safetyinstruction_attack#15 / 3165.52↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#13 / 3161.82↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#5 / 3180.85↑ higherSource ↗official