← Models

Model profile

Yi 34B Chat

01.AIdeveloper
2023-11-23release date
#182 / 312overall rank
9eval lineages

Evidence summary

Yi 34B Chat has an estimated overall rank of #182; its 90% source-sensitivity interval is #82–#221. Its behavior-only rank is #184; company governance moves the combined estimate to #182. Published evidence spans 9 evals and 5 of 7 behavior components. Its strongest relative result is CMoralEval (familial_morality, #1 of 26); its weakest is COMPL-AI AI-Identity Disclosure (score, #14 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#60 / 800.536Source ↗official
ChiSafetyBenchharmful_response_rate#8 / 140.65Source ↗official
ChiSafetyBenchmcq_score#10 / 1268.54Source ↗official
CMoralEvalfamilial_morality#1 / 260.71Source ↗official
CMoralEvalinternet_ethics#1 / 260.69Source ↗official
CMoralEvalpersonal_morality#1 / 260.66Source ↗official
CMoralEvalprofessional_ethics#1 / 260.7Source ↗official
CMoralEvalsocial_morality#1 / 260.71Source ↗official
COMPL-AI AI-Identity Disclosurescore#14 / 140.3562Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#4 / 140.5829Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#3 / 130.5387Source ↗official
CRiskEvaldeception_willingness#3 / 1718.68Source ↗official
CRiskEvaldesire_for_resource#9 / 1732.56Source ↗official
CRiskEvalharmful_goal#10 / 1741.77Source ↗official
CRiskEvalimprovement_intent#11 / 1751.23Source ↗official
CRiskEvalmalicious_coordination#8 / 1710.33Source ↗official
CRiskEvalself_preservation#12 / 1739.99Source ↗official
CRiskEvalsituational_awareness#14 / 1767.23Source ↗official
S-Evalbase_en_overall#20 / 2239.3Source ↗official
SafeDialBenchaggression#3 / 187.26Source ↗official
SafeDialBenchethics#2 / 187.68Source ↗official
SafeDialBenchfairness#9 / 187.337Source ↗official
SafeDialBenchlegality#1 / 188.117Source ↗official
SafeDialBenchmorality#1 / 187.52Source ↗official
SafeDialBenchprivacy#1 / 187.88Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#9 / 3324.14Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#7 / 3327.36Source ↗official
SALAD-Benchattack_enhanced_malicious_use#11 / 3322.76Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#9 / 3326.81Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#10 / 3322.6Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#8 / 3323.81Source ↗official
SALAD-Benchbase_human_autonomy_integrity#21 / 3391.73Source ↗official
SALAD-Benchbase_information_safety_harms#21 / 3393.23Source ↗official
SALAD-Benchbase_malicious_use#21 / 3389.36Source ↗official
SALAD-Benchbase_misinformation_harms#26 / 3387.74Source ↗official
SALAD-Benchbase_representation_toxicity#26 / 3381.07Source ↗official
SALAD-Benchbase_socioeconomic_harms#17 / 3389.19Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#21 / 3325Source ↗official
SALAD-Benchmcq_information_safety_harms#18 / 3331.67Source ↗official
SALAD-Benchmcq_malicious_use#18 / 3327.76Source ↗official
SALAD-Benchmcq_misinformation_harms#20 / 3326.43Source ↗official
SALAD-Benchmcq_representation_toxicity#20 / 3326.98Source ↗official
SALAD-Benchmcq_socioeconomic_harms#18 / 3331.67Source ↗official
SuperCLUE Safetyinstruction_attack#3 / 3172.41Source ↗official
SuperCLUE Safetyresponsible_ai#7 / 3169.09Source ↗official
SuperCLUE Safetytraditional_safety#10 / 3178.72Source ↗official