← Models

Model profile

Gemma 3 27B It

Googledeveloper
2025-03-12release date
#213 / 267overall rank
10eval lineages

Evidence summary

Gemma 3 27B It has an estimated overall rank of #213; its 90% source-sensitivity interval is #56–#245. Its behavior-only rank is #224; company governance moves the combined estimate to #213. Published evidence spans 10 evals and 6 of 7 behavior components. Its strongest relative result is HUMAINE Trust, Ethics and Safety (trust_ethics_safety_score, #17 of 54); its weakest is Confabulations (confabulation_rate, #52 of 52).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#250 / 3110.8954↓ lowerSource ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#24 / 3227↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#27 / 3238↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#27 / 3223↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#25 / 3242↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#30 / 3238↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#27 / 3235↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#28 / 3231↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#25 / 3224↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#19 / 3227↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#27 / 3230↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#28 / 3229↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#26 / 3233↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#103 / 1050.03↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#87 / 10569.44↓ lowerSource ↗official
Confabulationsconfabulation_rate#52 / 5266.34↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#117 / 26017.05↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#88 / 26090.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#169 / 26056.11↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#104 / 25897.36↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#17 / 5427.74↑ higherSource ↗official
PHAREbias_resistance_diagnostic#50 / 660.3801↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#67 / 700.603↑ higherSource ↗official
PHAREharm_resistance_diagnostic#45 / 700.9136↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#51 / 670.4007↑ higherSource ↗official
SOSBenchbiology_pvr#17 / 230.792↓ lowerSource ↗official
SOSBenchchemistry_pvr#19 / 230.646↓ lowerSource ↗official
SOSBenchmedicine_pvr#19 / 230.814↓ lowerSource ↗official
SOSBenchpharmacology_pvr#18 / 230.934↓ lowerSource ↗official
SOSBenchphysics_pvr#21 / 230.842↓ lowerSource ↗official
SOSBenchpsychology_pvr#18 / 230.792↓ lowerSource ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#8 / 256.3↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-14.2
Government46.5
Diplomacy62.8
Economy44.1
Society57.8