← Models

Model profile

Gemini 3.7 Flash

Googledeveloper
2026-08-13release date
#31 / 331Safety rank
#388 / 631Freedom rank

Evidence summary

Safety. Gemini 3.7 Flash has an estimated Safety rank of #31; its 90% source-sensitivity interval is #14–#97. Its behavior-only rank is #35; company governance moves the combined estimate to #31. Published Safety evidence spans 17 eval lineages and 6 of 7 components. Its strongest relative result is Manager Coercion Bench (fabrication_rate, #1 of 15); its weakest is SpeciEval (speciesism, #111 of 122).

Freedom. Gemini 3.7 Flash has an estimated Freedom rank of #388; its 90% source-sensitivity interval is #157–#535. Published Freedom evidence spans 2 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (overfit, #7 of 84); its weakest is SM-Bench (adversarial, #69 of 84).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#103 / 3380.6453Source ↗official
BullshitBench v2clear_pushback_rate#54 / 1140.33Source ↗official
CAIS Risk Indexagent_red_teaming#4 / 4938.5Source ↗official
CAIS Risk Indexbioweapons_assistance#27 / 5465.9Source ↗official
CAIS Risk Indexhle_overconfidence#26 / 5551.9Source ↗official
CAIS Risk Indexmachiavelli#21 / 5185.5Source ↗official
CAIS Risk Indexmask#43 / 5746.3Source ↗official
CAIS Risk Indexpolitical_manipulation#19 / 4845.1Source ↗official
CAIS Risk Indextextquests_harm#49 / 5424.4Source ↗official
Claude Fable 5.1 card — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct#5 / 119.2Source ↗official
Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct#7 / 159.2Source ↗official
Manager Coercion Benchcoercion_ladder_depth#20 / 358.7Source ↗official
Manager Coercion Benchfabrication_rate#1 / 150Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#10 / 227Source ↗official
Pander Scoreconversational_absolute_pander_score#12 / 2215.68Source ↗official
Pander Scoreinstructional_absolute_pander_score#15 / 2265.05Source ↗official
SM-Benchadversarial#13 / 8486.83Source ↗official
SM-Benchambiguous_interpretation#15 / 8490.48Source ↗official
SM-Benchanti_hallucination#23 / 8497.38Source ↗official
SM-Bencheq_boundaries#39 / 8465.73Source ↗official
SM-Benchoverfit#7 / 8493.99Source ↗official
SpeciEvalbelief_animal_sentience#44 / 1226.87Source ↗official
SpeciEvalland_animal_4ns#13 / 1224.17Source ↗official
SpeciEvalsea_animal_4ns#50 / 1224.67Source ↗official
SpeciEvalspeciesism#111 / 1222.77Source ↗official
TACbase_welfare_rate#46 / 8425.64Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
CAIS Risk Indexbioweapons_assistance#28 / 5465.9Source ↗official
SM-Benchadversarial#69 / 8486.83Source ↗official
SM-Bencheq_boundaries#39 / 8465.73Source ↗official
SM-Benchoverfit#7 / 8493.99Source ↗official