← Models

Model profile

Qwen2 7B Instruct

Alibabadeveloper
2024-06-07release date
#255 / 309overall rank
5eval lineages

Evidence summary

Qwen2 7B Instruct has an estimated overall rank of #255; its 90% source-sensitivity interval is #137–#294. Its behavior-only rank is #251; company governance moves the combined estimate to #255. Published evidence spans 5 evals and 4 of 7 behavior components. Its strongest relative result is PandaBench JBB direct-request panel (safety_rate, #14 of 46); its weakest is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #232 of 239).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Contextual MoralChoicehuman_agreement#16 / 220.35Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#85 / 24119.64Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#178 / 24183.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#134 / 24162.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#232 / 23966.32Source ↗official
IndoBias-Pairs — parity-aware culturally grounded biasparity_score#10 / 2686.94Source ↗official
M3-SafetyBenchoverall_score#9 / 1991.7Source ↗official
PandaBench JBB direct-request panelsafety_rate#14 / 460.99Source ↗official