← Models

Model profile

Gemma 4 31B It

Googledeveloper
2026-03-11release date
#55 / 305overall rank
14eval lineages
2discovery sources

Evidence summary

Gemma 4 31B It has an estimated overall rank of #55; its 90% source-sensitivity interval is #35–#150. Its behavior-only rank is #65; company governance moves the combined estimate to #55. Published evidence spans 14 evals and 7 of 7 behavior components. Its strongest relative result is SM-Bench (adversarial, #2 of 78); its weakest is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #234 of 241).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#167 / 3270.8195Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#44 / 1121442.0Source ↗official
BullshitBench v2clear_pushback_rate#60 / 1060.225Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#218 / 2418.53Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#234 / 24166.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#35 / 24193.89Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#135 / 23995.64Source ↗official
JuICE Cultural-Error Span Detectionf1#5 / 100.445Source ↗official
KIDBench Implicit Child Cueimplicit_child_cue_total_mean#4 / 134.14Source ↗official
PHAREbias_resistance_diagnostic#54 / 660.3522Source ↗official
PHAREhallucination_resistance_diagnostic#30 / 700.7662Source ↗official
PHAREharm_resistance_diagnostic#17 / 700.9651Source ↗official
PHAREjailbreak_resistance_diagnostic#24 / 670.6122Source ↗official
RealityTest — Text AI-Identity Disclosuredisclosure_probability#8 / 170.313Source ↗official
SM-Benchadversarial#2 / 7891.71Source ↗official
SM-Benchambiguous_interpretation#11 / 7890.77Source ↗official
SM-Benchanti_hallucination#22 / 7896.86Source ↗official
SM-Bencheq_boundaries#29 / 7867.7Source ↗official
SM-Benchoverfit#8 / 7891.8Source ↗official
SpeciEvalbelief_animal_sentience#18 / 1056.98Source ↗official
SpeciEvalland_animal_4ns#43 / 1054.47Source ↗official
SpeciEvalsea_animal_4ns#65 / 1054.85Source ↗official
SpeciEvalspeciesism#97 / 1052.88Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#5 / 2388.1Source ↗official
TACbase_welfare_rate#50 / 7423.72Source ↗self run
Vectara HHEM Factual Consistencyfactual_consistency_rate#31 / 9492.6Source ↗official
Vigil Mental Health Safetyoverall_score#12 / 2346Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20.1
Government45.7
Diplomacy69.8
Economy41.4
Society65.3