← Models

Model profile

Gemma 3 4B

Googledeveloper
#243 / 267overall rank
3eval lineages

Evidence summary

Gemma 3 4B has an estimated overall rank of #243; its 90% source-sensitivity interval is #195–#253. Its behavior-only rank is #259; company governance moves the combined estimate to #243. Published evidence spans 3 evals and 4 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #125 of 260); its weakest is AA-Omniscience (hallucination_rate, #311 of 311).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#311 / 3110.9883↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#76 / 10559.02↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#128 / 26016.02↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#125 / 26088.5↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#196 / 26048.89↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#166 / 25895.05↑ higherSource ↗official