← Models

Model profile

Chatglm3 6B

Z.aideveloper
2023-10-27release date
#232 / 309overall rank
10eval lineages

Evidence summary

Chatglm3 6B has an estimated overall rank of #232; its 90% source-sensitivity interval is #159–#257. Its behavior-only rank is #235; company governance moves the combined estimate to #232. Published evidence spans 10 evals and 4 of 7 behavior components. Its strongest relative result is SafeDialBench (ethics, #4 of 18); its weakest is ChiSafetyBench (mcq_score, #12 of 12).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChiSafetyBenchharmful_response_rate#12 / 141.08Source ↗official
ChiSafetyBenchmcq_score#12 / 1241.16Source ↗official
CMoralEvalfamilial_morality#11 / 260.43Source ↗official
CMoralEvalinternet_ethics#12 / 260.4Source ↗official
CMoralEvalpersonal_morality#11 / 260.42Source ↗official
CMoralEvalprofessional_ethics#11 / 260.43Source ↗official
CMoralEvalsocial_morality#11 / 260.43Source ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#9 / 1445.33Source ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#9 / 1494.67Source ↗official
FinEval Financial Security Knowledgefinancial_security_accuracy_pct#17 / 1948.2Source ↗official
FLAMESdata_protection#10 / 1338.16Source ↗official
FLAMESfairness#7 / 1337.75Source ↗official
FLAMESlegality#12 / 1328.26Source ↗official
FLAMESmorality#6 / 1344.78Source ↗official
FLAMESsafety#7 / 1332.63Source ↗official
S-Evalbase_en_overall#12 / 2257.7Source ↗official
SafeDialBenchaggression#15 / 187.017Source ↗official
SafeDialBenchethics#4 / 187.637Source ↗official
SafeDialBenchfairness#13 / 187.187Source ↗official
SafeDialBenchlegality#6 / 187.983Source ↗official
SafeDialBenchmorality#9 / 187.243Source ↗official
SafeDialBenchprivacy#9 / 187.567Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#18 / 3312.72Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#17 / 3312.05Source ↗official
SALAD-Benchattack_enhanced_malicious_use#17 / 3311.01Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#17 / 3313.16Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#17 / 3312.57Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#14 / 3317.75Source ↗official
SALAD-Benchbase_human_autonomy_integrity#20 / 3392.55Source ↗official
SALAD-Benchbase_information_safety_harms#23 / 3392.21Source ↗official
SALAD-Benchbase_malicious_use#20 / 3391.15Source ↗official
SALAD-Benchbase_misinformation_harms#22 / 3391.38Source ↗official
SALAD-Benchbase_representation_toxicity#19 / 3388.73Source ↗official
SALAD-Benchbase_socioeconomic_harms#19 / 3386.96Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#25 / 3317.78Source ↗official
SALAD-Benchmcq_information_safety_harms#21 / 3320.56Source ↗official
SALAD-Benchmcq_malicious_use#22 / 3319.55Source ↗official
SALAD-Benchmcq_misinformation_harms#22 / 3320.95Source ↗official
SALAD-Benchmcq_representation_toxicity#24 / 3317.81Source ↗official
SALAD-Benchmcq_socioeconomic_harms#22 / 3322.78Source ↗official
SORRY-Benchavg#34 / 510.36Source ↗official
SuperCLUE Safetyinstruction_attack#9 / 3168.97Source ↗official
SuperCLUE Safetyresponsible_ai#17 / 3156.36Source ↗official
SuperCLUE Safetytraditional_safety#28 / 3165.96Source ↗official