← Models

Model profile

Internlm2 Chat 20B

InternLMdeveloper
2024-01-17release date
#59 / 267overall rank
4eval lineages

Evidence summary

Internlm2 Chat 20B has an estimated overall rank of #59; its 90% source-sensitivity interval is #21–#222. Published evidence spans 4 evals and 5 of 7 behavior components. Its strongest relative result is SALAD-Bench (base_representation_toxicity, #2 of 33); its weakest is ChineseSafe (score, #15 of 22).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
ChineseSafescore#15 / 2253.67↑ higherSource ↗official
CMoralEvalfamilial_morality#3 / 260.56↑ higherSource ↗official
CMoralEvalinternet_ethics#3 / 260.54↑ higherSource ↗official
CMoralEvalpersonal_morality#3 / 260.52↑ higherSource ↗official
CMoralEvalprofessional_ethics#3 / 260.54↑ higherSource ↗official
CMoralEvalsocial_morality#3 / 260.54↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#88 / 26020.41↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#41 / 26093.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#128 / 26071.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#127 / 25896.41↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#15 / 3315.52↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#21 / 337.17↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#15 / 3312.56↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#19 / 339.38↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#20 / 3310.12↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#17 / 3312.55↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#3 / 3398.54↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#11 / 3396.68↑ higherSource ↗official
SALAD-Benchbase_malicious_use#2 / 3398.88↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#2 / 3398.67↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#2 / 3397.53↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#3 / 3395.77↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#8 / 3361.11↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#7 / 3351.11↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#6 / 3357.24↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#6 / 3359.05↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#6 / 3357.92↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#10 / 3346.11↑ higherSource ↗official