← Models

Model profile

Gemini 3.5 Flash

Googledeveloper
2026-05-19release date
#45 / 309overall rank
23eval lineages

Evidence summary

Gemini 3.5 Flash has an estimated overall rank of #45; its 90% source-sensitivity interval is #26–#191. Its behavior-only rank is #55; company governance moves the combined estimate to #45. Published evidence spans 23 evals and 7 of 7 behavior components. Its strongest relative result is SM-Bench (overfit, #1 of 79); its weakest is SpeciEval (speciesism, #112 of 113).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#80 / 3280.6176Source ↗official
ANIMAscore#11 / 220.7004Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#51 / 1121440.0Source ↗official
BioSecBench-Refusalbalanced_refusal_score#1 / 100.5041Source ↗official
BullshitBench v2clear_pushback_rate#66 / 1060.195Source ↗official
CAIS Risk Indexagent_red_teaming#21 / 4472.3Source ↗official
CAIS Risk Indexbioweapons_assistance#32 / 4978.3Source ↗official
CAIS Risk Indexhle_overconfidence#22 / 5053.5Source ↗official
CAIS Risk Indexmachiavelli#31 / 4692.3Source ↗official
CAIS Risk Indexmask#47 / 5257.4Source ↗official
CAIS Risk Indexpolitical_manipulation#26 / 3353.4Source ↗official
CAIS Risk Indextextquests_harm#33 / 4920.7Source ↗official
DystopiaBenchbasaglia_score#45 / 5070.93Source ↗official
DystopiaBenchbaudrillard_score#39 / 5066.63Source ↗official
DystopiaBenchhuxley_score#43 / 5079.23Source ↗official
DystopiaBenchlaguardia_score#44 / 5070.9Source ↗official
DystopiaBenchorwell_score#45 / 5075.17Source ↗official
DystopiaBenchpetrov_score#26 / 5072.97Source ↗official
Governance Decay under Passive Context Compactiongovernance_retention_score#2 / 796Source ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#12 / 1360.5Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#2 / 5429.37Source ↗official
MACHIAVELLIdeception_relative_random_pct#32 / 5092.3Source ↗official
MORUscore#5 / 1377.89Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#8 / 194Source ↗official
PHAREhallucination_resistance_diagnostic#12 / 700.8218Source ↗official
PHAREharm_resistance_diagnostic#19 / 700.9594Source ↗official
PHAREjailbreak_resistance_diagnostic#46 / 670.4264Source ↗official
SM-Benchadversarial#1 / 7992.2Source ↗official
SM-Benchambiguous_interpretation#37 / 7986.01Source ↗official
SM-Benchanti_hallucination#51 / 7988.48Source ↗official
SM-Bencheq_boundaries#25 / 7968.54Source ↗official
SM-Benchoverfit#1 / 7998.36Source ↗official
SpeciEvalbelief_animal_sentience#72 / 1136.68Source ↗official
SpeciEvalland_animal_4ns#3 / 1133.62Source ↗official
SpeciEvalsea_animal_4ns#36 / 1134.65Source ↗official
SpeciEvalspeciesism#112 / 1133.88Source ↗official
TACbase_welfare_rate#69 / 7617.31Source ↗official
ToolPrivacyBenchprivate_mt_poi#1 / 919.19Source ↗official
ToolPrivacyBenchpublic_mt_poi#8 / 919.86Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-13.4
Government48.5
Diplomacy61.4
Economy47.1
Society58