← Models

Model profile

Gemini 3.1 Flash Lite

Googledeveloper
2026-03-03release date
#125 / 267overall rank
20eval lineages

Evidence summary

Gemini 3.1 Flash Lite has an estimated overall rank of #125; its 90% source-sensitivity interval is #55–#186. Its behavior-only rank is #140; company governance moves the combined estimate to #125. Published evidence spans 20 evals and 7 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 102); its weakest is MANTA (AWVS, #7 of 7).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#168 / 3110.8163↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#84 / 1050.11↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#7 / 4345.4↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#39 / 4889.8↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#39 / 4972.8↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#40 / 4595.8↓ lowerSource ↗official
CAIS Risk Indexmask#45 / 5155.8↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#23 / 3251.9↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#14 / 4815.6↓ lowerSource ↗official
DystopiaBenchbasaglia_score#40 / 5069.4↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#35 / 5065.2↓ lowerSource ↗official
DystopiaBenchhuxley_score#44 / 5080.73↓ lowerSource ↗official
DystopiaBenchlaguardia_score#41 / 5069.77↓ lowerSource ↗official
DystopiaBenchorwell_score#39 / 5073.97↓ lowerSource ↗official
DystopiaBenchpetrov_score#40 / 5078.43↓ lowerSource ↗official
FORTRESSaverage_risk_score#36 / 4951.05↓ lowerSource ↗official
FORTRESSover_refusal_score#11 / 462.15↓ lowerSource ↗official
MACHIAVELLIdeception_relative_random_pct#42 / 5095.8↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#10 / 317.633↓ lowerSource ↗self run
MANTAAWMS#5 / 70.401↑ higherSource ↗official
MANTAAWVS#7 / 70.309↑ higherSource ↗official
MASKlying_probability_pct#45 / 5351.6↓ lowerSource ↗official
PHAREbias_resistance_diagnostic#37 / 660.4457↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#19 / 700.8057↑ higherSource ↗official
PHAREharm_resistance_diagnostic#13 / 700.9694↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#37 / 670.4772↑ higherSource ↗official
RefusalBenchyouden_j#10 / 190.4533↑ higherSource ↗official
SM-Benchadversarial#2 / 7391.71↑ higherSource ↗official
SM-Benchambiguous_interpretation#45 / 7382.44↑ higherSource ↗official
SM-Benchanti_hallucination#27 / 7394.76↑ higherSource ↗official
SM-Bencheq_boundaries#30 / 7366.85↑ higherSource ↗official
SM-Benchoverfit#2 / 7397.81↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#1 / 1027↑ higherSource ↗official
SpeciEvalland_animal_4ns#41 / 1024.46↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#49 / 1024.75↓ lowerSource ↗official
SpeciEvalspeciesism#99 / 1023.225↓ lowerSource ↗official
TACbase_welfare_rate#26 / 6830.13↑ higherSource ↗self run
VETO Misfired Alignmentmisfired_alignment_rate_pct#13 / 258.3↓ lowerSource ↗official
Vigil Mental Health Safetyoverall_score#13 / 2338↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression3.82
Traditional ↔ Secular0.967