← Models

Model profile

Gemini 3.8 Flash

Googledeveloper
2026-09-02release date
#10 / 331Safety rank
#443 / 631Freedom rank

Evidence summary

Safety. Gemini 3.8 Flash has an estimated Safety rank of #10; its 90% source-sensitivity interval is #7–#63. Its behavior-only rank is #12; company governance moves the combined estimate to #10. Published Safety evidence spans 12 eval lineages and 6 of 7 components. Its strongest relative result is CAIS Risk Index (agent_red_teaming, #2 of 49); its weakest is CAIS Risk Index (textquests_harm, #49 of 54).

Freedom. Gemini 3.8 Flash has an estimated Freedom rank of #443; its 90% source-sensitivity interval is #278–#554. Published Freedom evidence spans 2 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (overfit, #4 of 84); its weakest is SM-Bench (adversarial, #74 of 84).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#78 / 3380.519Source ↗official
CAIS Risk Indexagent_red_teaming#2 / 4929.8Source ↗official
CAIS Risk Indexbioweapons_assistance#25 / 5465.6Source ↗official
CAIS Risk Indexhle_overconfidence#25 / 5551.3Source ↗official
CAIS Risk Indexmachiavelli#23 / 5186.5Source ↗official
CAIS Risk Indexmask#39 / 5741.4Source ↗official
CAIS Risk Indexpolitical_manipulation#22 / 4846.3Source ↗official
CAIS Risk Indextextquests_harm#49 / 5424.4Source ↗official
Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct#2 / 155.5Source ↗official
SM-Benchadversarial#9 / 8488.29Source ↗official
SM-Benchambiguous_interpretation#37 / 8486.9Source ↗official
SM-Benchanti_hallucination#29 / 8496.34Source ↗official
SM-Bencheq_boundaries#50 / 8462.36Source ↗official
SM-Benchoverfit#4 / 8495.08Source ↗official
SpeciEvalbelief_animal_sentience#53 / 1226.83Source ↗official
SpeciEvalland_animal_4ns#33 / 1224.33Source ↗official
SpeciEvalsea_animal_4ns#61 / 1224.72Source ↗official
SpeciEvalspeciesism#101 / 1222.58Source ↗official
TACbase_welfare_rate#28 / 8430.77Source ↗self run

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
CAIS Risk Indexbioweapons_assistance#30 / 5465.6Source ↗official
SM-Benchadversarial#74 / 8488.29Source ↗official
SM-Bencheq_boundaries#50 / 8462.36Source ↗official
SM-Benchoverfit#4 / 8495.08Source ↗official