← Models

Model profile

Chatglm3 6B

Z.aideveloper
2023-10-27release date
#202 / 267overall rank
9eval lineages

Evidence summary

Chatglm3 6B has an estimated overall rank of #202; its 90% source-sensitivity interval is #127–#222. Published evidence spans 9 evals and 4 of 7 behavior components. Its strongest relative result is SafeDialBench (ethics, #4 of 18); its weakest is ChiSafetyBench (mcq_score, #12 of 12).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
ChiSafetyBenchharmful_response_rate#12 / 141.08↓ lowerSource ↗official
ChiSafetyBenchmcq_score#12 / 1241.16↑ higherSource ↗official
CMoralEvalfamilial_morality#11 / 260.43↑ higherSource ↗official
CMoralEvalinternet_ethics#12 / 260.4↑ higherSource ↗official
CMoralEvalpersonal_morality#11 / 260.42↑ higherSource ↗official
CMoralEvalprofessional_ethics#11 / 260.43↑ higherSource ↗official
CMoralEvalsocial_morality#11 / 260.43↑ higherSource ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#9 / 1445.33↑ higherSource ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#9 / 1494.67↑ higherSource ↗official
FLAMESdata_protection#10 / 1338.16↑ higherSource ↗official
FLAMESfairness#7 / 1337.75↑ higherSource ↗official
FLAMESlegality#12 / 1328.26↑ higherSource ↗official
FLAMESmorality#6 / 1344.78↑ higherSource ↗official
FLAMESsafety#7 / 1332.63↑ higherSource ↗official
S-Evalbase_en_overall#12 / 2257.7↑ higherSource ↗official
SafeDialBenchaggression#15 / 187.017↑ higherSource ↗official
SafeDialBenchethics#4 / 187.637↑ higherSource ↗official
SafeDialBenchfairness#13 / 187.187↑ higherSource ↗official
SafeDialBenchlegality#6 / 187.983↑ higherSource ↗official
SafeDialBenchmorality#9 / 187.243↑ higherSource ↗official
SafeDialBenchprivacy#9 / 187.567↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#18 / 3312.72↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#17 / 3312.05↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#17 / 3311.01↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#17 / 3313.16↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#17 / 3312.57↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#14 / 3317.75↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#20 / 3392.55↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#23 / 3392.21↑ higherSource ↗official
SALAD-Benchbase_malicious_use#20 / 3391.15↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#22 / 3391.38↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#19 / 3388.73↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#19 / 3386.96↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#25 / 3317.78↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#21 / 3320.56↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#22 / 3319.55↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#22 / 3320.95↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#24 / 3317.81↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#22 / 3322.78↑ higherSource ↗official
SORRY-Benchavg#34 / 510.36↓ lowerSource ↗official
SuperCLUE Safetyinstruction_attack#9 / 3168.97↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#17 / 3156.36↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#28 / 3165.96↑ higherSource ↗official