← Models

Model profile

Internlm2 20B Sft

2024-01-11release date
1eval lineages

Evidence summary

Published evidence spans 1 evals and 3 of 7 behavior components. Its strongest relative result is SafeDialBench (legality, #7 of 18); its weakest is SafeDialBench (fairness, #18 of 18).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
SafeDialBenchaggression#17 / 186.957↑ higherSource ↗official
SafeDialBenchethics#11 / 187.547↑ higherSource ↗official
SafeDialBenchfairness#18 / 187.1↑ higherSource ↗official
SafeDialBenchlegality#7 / 187.977↑ higherSource ↗official
SafeDialBenchmorality#11 / 187.213↑ higherSource ↗official
SafeDialBenchprivacy#7 / 187.597↑ higherSource ↗official