← Models

Model profile

Internlm Chat 20B

InternLMdeveloper
2023-09-18release date
#67 / 267overall rank
4eval lineages

Evidence summary

Internlm Chat 20B has an estimated overall rank of #67; its 90% source-sensitivity interval is #20–#185. Its behavior-only rank is #63; company governance moves the combined estimate to #67. Published evidence spans 4 evals and 4 of 7 behavior components. Its strongest relative result is SALAD-Bench (base_information_safety_harms, #1 of 33); its weakest is SALAD-Bench (mcq_representation_toxicity, #29 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Fake Alignment (FINE)multiple_choice_safe_decision_rate#3 / 1469.33↑ higherSource ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#7 / 1496↑ higherSource ↗official
FLAMESdata_protection#1 / 1363.16↑ higherSource ↗official
FLAMESfairness#1 / 1352.61↑ higherSource ↗official
FLAMESlegality#2 / 1371.74↑ higherSource ↗official
FLAMESmorality#1 / 1354.23↑ higherSource ↗official
FLAMESsafety#3 / 1351.05↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#7 / 3329.74↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#9 / 3323.45↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#7 / 3329.36↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#8 / 3327.96↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#7 / 3331.79↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#7 / 3327.27↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#2 / 3398.89↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#1 / 3399.8↑ higherSource ↗official
SALAD-Benchbase_malicious_use#5 / 3398.44↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#6 / 3397↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#8 / 3393.39↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#2 / 3396.36↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#29 / 333.333↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#27 / 334.167↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#29 / 333.974↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#29 / 333.333↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#29 / 333.229↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#27 / 337.778↑ higherSource ↗official
SuperCLUE Safetyinstruction_attack#9 / 3168.97↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#11 / 3165.45↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#25 / 3168.09↑ higherSource ↗official