← Models

Model profile

Gemma 7B IT

Googledeveloper
2024-02-21release date
#213 / 333Safety rank
#312 / 645Freedom rank

Evidence summary

Safety. Gemma 7B IT has an estimated Safety rank of #213; its 90% source-sensitivity interval is #93–#262. Its behavior-only rank is #228; company governance moves the combined estimate to #213. Published Safety evidence spans 7 eval lineages and 7 of 7 components. Its strongest relative result is SALAD-Bench (base_representation_toxicity, #6 of 33); its weakest is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #231 of 246).

Freedom. Gemma 7B IT has an estimated Freedom rank of #312; its 90% source-sensitivity interval is #215–#416. Published Freedom evidence spans 6 eval lineages and 1 of 1 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #15 of 246); its weakest is SALAD-Bench (base_representation_toxicity, #28 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#86 / 24820.41Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#206 / 24878.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#173 / 24854.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#231 / 24677.05Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#32 / 391.314Source ↗official
Microsoft Phi Safety Panelsharmful_continuation#6 / 100.013Source ↗official
Microsoft Phi Safety Panelsharmful_summarization#3 / 100.103Source ↗official
Microsoft Phi Safety Panelsjailbreak#4 / 100.114Source ↗official
Microsoft Phi Safety Panelsthird_party_harm#8 / 100.383Source ↗official
OR-Benchover_refusal_rate#7 / 2526.3Source ↗official
OR-Benchtoxic_acceptance_rate#17 / 2514.5Source ↗official
S-Evalbase_en_overall#10 / 2261.8Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#16 / 3313.36Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#11 / 3322.8Source ↗official
SALAD-Benchattack_enhanced_malicious_use#19 / 339.95Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#12 / 3318.91Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#15 / 3317.56Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#18 / 3312.12Source ↗official
SALAD-Benchbase_human_autonomy_integrity#17 / 3394.82Source ↗official
SALAD-Benchbase_information_safety_harms#9 / 3397.49Source ↗official
SALAD-Benchbase_malicious_use#17 / 3393.54Source ↗official
SALAD-Benchbase_misinformation_harms#12 / 3395.57Source ↗official
SALAD-Benchbase_representation_toxicity#6 / 3394.42Source ↗official
SALAD-Benchbase_socioeconomic_harms#22 / 3386.13Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#15 / 3340.56Source ↗official
SALAD-Benchmcq_information_safety_harms#12 / 3340.56Source ↗official
SALAD-Benchmcq_malicious_use#14 / 3340.38Source ↗official
SALAD-Benchmcq_misinformation_harms#16 / 3338.1Source ↗official
SALAD-Benchmcq_representation_toxicity#16 / 3338.85Source ↗official
SALAD-Benchmcq_socioeconomic_harms#14 / 3338.89Source ↗official
SORRY-Benchavg#16 / 510.18Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#43 / 24878.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#73 / 24854.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#15 / 24677.05Source ↗official
Microsoft Phi Safety Panelsharmful_continuation#4 / 100.013Source ↗official
Microsoft Phi Safety Panelsharmful_summarization#8 / 100.103Source ↗official
Microsoft Phi Safety Panelsjailbreak#7 / 100.114Source ↗official
Microsoft Phi Safety Panelsthird_party_harm#3 / 100.383Source ↗official
OR-Benchover_refusal_rate#7 / 2526.3Source ↗official
OR-Benchtoxic_acceptance_rate#9 / 2514.5Source ↗official
S-Evalbase_en_overall#13 / 2261.8Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#17 / 3313.36Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#23 / 3322.8Source ↗official
SALAD-Benchattack_enhanced_malicious_use#15 / 339.95Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#22 / 3318.91Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#19 / 3317.56Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#15 / 3312.12Source ↗official
SALAD-Benchbase_human_autonomy_integrity#17 / 3394.82Source ↗official
SALAD-Benchbase_information_safety_harms#25 / 3397.49Source ↗official
SALAD-Benchbase_malicious_use#17 / 3393.54Source ↗official
SALAD-Benchbase_misinformation_harms#22 / 3395.57Source ↗official
SALAD-Benchbase_representation_toxicity#28 / 3394.42Source ↗official
SALAD-Benchbase_socioeconomic_harms#12 / 3386.13Source ↗official
SORRY-Benchavg#36 / 510.18Source ↗official