← Models

Model profile

Gemma 3 27B It

Googledeveloper
2025-03-12release date
#235 / 312overall rank
13eval lineages

Evidence summary

Gemma 3 27B It has an estimated overall rank of #235; its 90% source-sensitivity interval is #56–#274. Its behavior-only rank is #248; company governance moves the combined estimate to #235. Published evidence spans 13 evals and 6 of 7 behavior components. Its strongest relative result is HUMAINE Trust, Ethics and Safety (trust_ethics_safety_score, #17 of 54); its weakest is Confabulations (confabulation_rate, #52 of 52).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#282 / 3300.9209Source ↗official
AgentDrive Safety Compliancescr#33 / 4866.25Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#24 / 3227Source ↗official
AILuminate General Purpose AI Chatdefamation#27 / 3238Source ↗official
AILuminate General Purpose AI Chathate#27 / 3223Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#25 / 3242Source ↗official
AILuminate General Purpose AI Chatintellectual_property#30 / 3238Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#27 / 3235Source ↗official
AILuminate General Purpose AI Chatprivacy#28 / 3231Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#25 / 3224Source ↗official
AILuminate General Purpose AI Chatsexual_content#19 / 3227Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#27 / 3230Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#28 / 3229Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#26 / 3233Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#40 / 1121445.0Source ↗official
BullshitBench v2clear_pushback_rate#104 / 1060.03Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#86 / 10469.44Source ↗official
Confabulationsconfabulation_rate#52 / 5266.34Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#106 / 24117.05Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#78 / 24190.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#153 / 24156.11Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#93 / 23997.36Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#17 / 5427.74Source ↗official
PHAREbias_resistance_diagnostic#50 / 660.3801Source ↗official
PHAREhallucination_resistance_diagnostic#67 / 700.603Source ↗official
PHAREharm_resistance_diagnostic#45 / 700.9136Source ↗official
PHAREjailbreak_resistance_diagnostic#51 / 670.4007Source ↗official
SOSBenchbiology_pvr#17 / 230.792Source ↗official
SOSBenchchemistry_pvr#19 / 230.646Source ↗official
SOSBenchmedicine_pvr#19 / 230.814Source ↗official
SOSBenchpharmacology_pvr#18 / 230.934Source ↗official
SOSBenchphysics_pvr#21 / 230.842Source ↗official
SOSBenchpsychology_pvr#18 / 230.792Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#31 / 9492.6Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#8 / 256.3Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-14.2
Government46.5
Diplomacy62.8
Economy44.1
Society57.8

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions1.11
CCP-narrative alignment — China topics1.11
CCP-narrative alignment — non-China controls1.11