← Models

Model profile

Gemini 3.5 Flash

Googledeveloper
2026-05-19release date
#66 / 267overall rank
20eval lineages

Evidence summary

Gemini 3.5 Flash has an estimated overall rank of #66; its 90% source-sensitivity interval is #22–#173. Its behavior-only rank is #72; company governance moves the combined estimate to #66. Published evidence spans 20 evals and 7 of 7 behavior components. Its strongest relative result is SM-Bench (overfit, #1 of 73); its weakest is SpeciEval (speciesism, #101 of 102).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#70 / 3110.6028↓ lowerSource ↗official
ANIMAscore#10 / 190.7004↑ higherSource ↗official
BioSecBench-Refusalbalanced_refusal_score#1 / 100.5041↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#65 / 1050.195↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#20 / 4372.3↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#31 / 4878.3↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#21 / 4953.5↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#31 / 4592.3↓ lowerSource ↗official
CAIS Risk Indexmask#46 / 5157.4↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#25 / 3253.4↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#33 / 4820.7↓ lowerSource ↗official
DystopiaBenchbasaglia_score#45 / 5070.93↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#39 / 5066.63↓ lowerSource ↗official
DystopiaBenchhuxley_score#43 / 5079.23↓ lowerSource ↗official
DystopiaBenchlaguardia_score#44 / 5070.9↓ lowerSource ↗official
DystopiaBenchorwell_score#45 / 5075.17↓ lowerSource ↗official
DystopiaBenchpetrov_score#26 / 5072.97↓ lowerSource ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#12 / 1360.5↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#2 / 5429.37↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#32 / 5092.3↓ lowerSource ↗official
MORUscore#5 / 1377.89↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#12 / 700.8218↑ higherSource ↗official
PHAREharm_resistance_diagnostic#19 / 700.9594↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#46 / 670.4264↑ higherSource ↗official
SM-Benchadversarial#1 / 7392.2↑ higherSource ↗official
SM-Benchambiguous_interpretation#32 / 7386.01↑ higherSource ↗official
SM-Benchanti_hallucination#46 / 7388.48↑ higherSource ↗official
SM-Bencheq_boundaries#23 / 7368.54↑ higherSource ↗official
SM-Benchoverfit#1 / 7398.36↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#64 / 1026.68↑ higherSource ↗official
SpeciEvalland_animal_4ns#3 / 1023.62↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#29 / 1024.65↓ lowerSource ↗official
SpeciEvalspeciesism#101 / 1023.88↓ lowerSource ↗official
TACbase_welfare_rate#63 / 6817.31↑ higherSource ↗official
ToolPrivacyBenchprivate_mt_poi#1 / 919.19↓ lowerSource ↗official
ToolPrivacyBenchpublic_mt_poi#8 / 919.86↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-13.4
Government48.5
Diplomacy61.4
Economy47.1
Society58