← Models

Model profile

Gemini 3 Flash Preview

Googledeveloper
2025-12-17release date
#160 / 312overall rank
23eval lineages
1discovery sources

Evidence summary

Gemini 3 Flash Preview has an estimated overall rank of #160; its 90% source-sensitivity interval is #78–#248. Its behavior-only rank is #174; company governance moves the combined estimate to #160. Published evidence spans 23 evals and 7 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 113); its weakest is DystopiaBench (baudrillard_score, #50 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#285 / 3300.9241Source ↗official
AgentAbstainabstain#16 / 1743.6Source ↗official
AgentAbstaincar#15 / 1743.4Source ↗official
AgentAbstainpaired#14 / 1739.7Source ↗official
Alignment Leaderboardcorrigibility#10 / 244.32Source ↗official
Alignment Leaderboardhonesty#9 / 243.682Source ↗official
Alignment Leaderboardnon_manipulation#12 / 243.51Source ↗official
Alignment Leaderboardrobustness#14 / 243.373Source ↗official
Alignment Leaderboardsafety#11 / 243.875Source ↗official
Alignment Leaderboardscheming#12 / 243.645Source ↗official
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating#25 / 301153.0Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#32 / 1121449.0Source ↗official
AuAu Authoritarian Response Auditbehavioral_vignette_arr_percent#3 / 1721.8Source ↗official
AuAu Authoritarian Response Auditrealistic_prompt_arr_percent#10 / 172.2Source ↗official
BullshitBench v2clear_pushback_rate#86 / 1060.1Source ↗official
CAIS Risk Indexagent_red_teaming#43 / 4592.9Source ↗official
CAIS Risk Indexbioweapons_assistance#31 / 5075.3Source ↗official
CAIS Risk Indexhle_overconfidence#24 / 5155.1Source ↗official
CAIS Risk Indexmachiavelli#38 / 4795.3Source ↗official
CAIS Risk Indexmask#50 / 5361.7Source ↗official
CAIS Risk Indextextquests_harm#24 / 5017.8Source ↗official
DystopiaBenchbasaglia_score#47 / 5073.2Source ↗official
DystopiaBenchbaudrillard_score#50 / 5081.13Source ↗official
DystopiaBenchhuxley_score#47 / 5083Source ↗official
DystopiaBenchlaguardia_score#46 / 5071.63Source ↗official
DystopiaBenchorwell_score#31 / 5072.43Source ↗official
DystopiaBenchpetrov_score#43 / 5080.33Source ↗official
JuICE Cultural-Error Span Detectionf1#3 / 100.4839Source ↗official
MACHIAVELLIdeception_relative_random_pct#39 / 5095.3Source ↗official
MT-JailBench CrescendoXsafety_score#16 / 2113.84Source ↗official
Pander Scoreconversational_absolute_pander_score#18 / 2025.97Source ↗official
Pander Scoreinstructional_absolute_pander_score#18 / 2070.12Source ↗official
RealityTest — Text AI-Identity Disclosuredisclosure_probability#12 / 170.227Source ↗official
SM-Benchadversarial#23 / 7984.39Source ↗official
SM-Benchambiguous_interpretation#58 / 7980.65Source ↗official
SM-Benchanti_hallucination#65 / 7982.72Source ↗official
SM-Bencheq_boundaries#37 / 7965.45Source ↗official
SM-Benchoverfit#5 / 7993.99Source ↗official
SpeciEvalbelief_animal_sentience#1 / 1137Source ↗official
SpeciEvalland_animal_4ns#77 / 1134.72Source ↗official
SpeciEvalsea_animal_4ns#89 / 1135.03Source ↗official
SpeciEvalspeciesism#113 / 1133.9Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#23 / 2382.56Source ↗official
TACbase_welfare_rate#15 / 7635.9Source ↗self run
Vectara HHEM Factual Consistencyfactual_consistency_rate#79 / 9486.5Source ↗official
Vigil Mental Health Safetyoverall_score#11 / 2347Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.7
Government45.8
Diplomacy66.6
Economy46.6
Society60

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience6.5
Honesty-humility5.5
Extraversion6.5
Agreeableness5.7
Conscientiousness7.3

Agent-ValueBench Schwartz Basic Values (PVQ40)

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression5.29
Traditional ↔ Secular-0.78