← Models

Model profile

Gemini 3 Flash Preview

Googledeveloper
2025-12-17release date
#146 / 267overall rank
15eval lineages
1discovery sources

Evidence summary

Gemini 3 Flash Preview has an estimated overall rank of #146; its 90% source-sensitivity interval is #69–#204. Its behavior-only rank is #154; company governance moves the combined estimate to #146. Published evidence spans 15 evals and 6 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 102); its weakest is DystopiaBench (baudrillard_score, #50 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#255 / 3110.9018↓ lowerSource ↗official
AgentAbstainabstain#16 / 1743.6↑ higherSource ↗official
AgentAbstaincar#15 / 1743.4↑ higherSource ↗official
AgentAbstainpaired#14 / 1739.7↑ higherSource ↗official
Alignment Leaderboardcorrigibility#10 / 244.32↑ higherSource ↗official
Alignment Leaderboardhonesty#9 / 243.682↑ higherSource ↗official
Alignment Leaderboardnon_manipulation#12 / 243.51↑ higherSource ↗official
Alignment Leaderboardrobustness#14 / 243.373↑ higherSource ↗official
Alignment Leaderboardsafety#11 / 243.875↑ higherSource ↗official
Alignment Leaderboardscheming#12 / 243.645↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#85 / 1050.1↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#41 / 4392.9↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#29 / 4875.3↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#22 / 4955.1↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#37 / 4595.3↓ lowerSource ↗official
CAIS Risk Indexmask#48 / 5161.7↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#24 / 4817.8↓ lowerSource ↗official
DystopiaBenchbasaglia_score#47 / 5073.2↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#50 / 5081.13↓ lowerSource ↗official
DystopiaBenchhuxley_score#47 / 5083↓ lowerSource ↗official
DystopiaBenchlaguardia_score#46 / 5071.63↓ lowerSource ↗official
DystopiaBenchorwell_score#31 / 5072.43↓ lowerSource ↗official
DystopiaBenchpetrov_score#43 / 5080.33↓ lowerSource ↗official
MACHIAVELLIdeception_relative_random_pct#39 / 5095.3↓ lowerSource ↗official
SM-Benchadversarial#22 / 7384.39↑ higherSource ↗official
SM-Benchambiguous_interpretation#53 / 7380.65↑ higherSource ↗official
SM-Benchanti_hallucination#60 / 7382.72↑ higherSource ↗official
SM-Bencheq_boundaries#32 / 7365.45↑ higherSource ↗official
SM-Benchoverfit#5 / 7393.99↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#1 / 1027↑ higherSource ↗official
SpeciEvalland_animal_4ns#69 / 1024.72↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#79 / 1025.03↓ lowerSource ↗official
SpeciEvalspeciesism#102 / 1023.9↓ lowerSource ↗official
TACbase_welfare_rate#15 / 6835.9↑ higherSource ↗self run
Vigil Mental Health Safetyoverall_score#11 / 2347↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.7
Government45.8
Diplomacy66.6
Economy46.6
Society60

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience6.5
Honesty-humility5.5
Extraversion6.5
Agreeableness5.7
Conscientiousness7.3

Agent-ValueBench Schwartz Basic Values (PVQ40)

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression5.29
Traditional ↔ Secular-0.78