Developer
53 indexed models; 32 currently meet the evidence threshold for the overall ranking. Together they have results from 115 evaluations.
Company governance evidence
Google is represented at +0.81 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by Google
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Gemini 3.8 Flash | 15 | 6/7 | 17 | 2026-09-02 |
| Gemini 3.7 Flash | 18 | 6/7 | 19 | 2026-08-13 |
| Gemini 3.6 Flash | 10 | 6/7 | 20 | 2026-07-21 |
| Gemini 2.5 Pro Exp | 4 | 3/7 | 38 | 2025-03-25 |
| Gemini 3.1 Pro | 6 | 5/7 | 41 | 2026-02-19 |
| Gemini 3.1 Pro Preview | 34 | 7/7 | 63 | 2026-02-19 |
| Gemini 3.5 Flash | 30 | 7/7 | 65 | 2026-05-19 |
| Gemma 4 26B A4B IT | 8 | 5/7 | 72 | 2026-03-11 |
| Gemma 4 31B IT | 17 | 7/7 | 79 | 2026-03-11 |
| Gemini 1.0 Pro | 8 | 6/7 | 83 | 2023-12-06 |
| Gemini 1.5 Pro | 22 | 7/7 | 104 | 2024-02-15 |
| Gemini 3.5 Flash Lite | 9 | 5/7 | 108 | 2026-07-21 |
| Gemini 2.5 Pro | 46 | 7/7 | 111 | 2025-03-25 |
| Gemini 3 Pro Preview | 27 | 7/7 | 120 | 2025-11-18 |
| Gemini 1.5 Flash | 17 | 7/7 | 121 | 2024-05-14 |
| Gemma 4 E4B | 3 | 4/7 | 152 | 2026-03-31 |
| Gemini 2.5 Flash | 40 | 7/7 | 157 | 2025-04-17 |
| Gemini 2.0 Pro Preview | 6 | 3/7 | 166 | 2025-02-05 |
| Gemini 3.1 Flash Lite | 27 | 7/7 | 179 | 2026-03-03 |
| Gemini 2.0 Flash | 20 | 7/7 | 197 | 2024-12-11 |
| Gemma 2B IT | 5 | 6/7 | 204 | 2024-02-21 |
| Gemini 3 Flash Preview | 29 | 7/7 | 212 | 2025-12-17 |
| Gemma 7B IT | 7 | 7/7 | 213 | 2024-02-21 |
| Gemini 2.0 Flash Lite | 5 | 5/7 | 219 | 2025-02-25 |
| Gemma 2 9B IT | 9 | 7/7 | 227 | 2024-06-27 |
| Gemini 2.5 Flash Lite | 23 | 7/7 | 228 | 2025-06-17 |
| Gemma 3 12B | 9 | 6/7 | 229 | 2025-03-10 |
| Gemini 2.0 Flash Lite Preview | 7 | 5/7 | 236 | 2025-02-05 |
| Gemma 2 27B IT | 6 | 7/7 | 242 | 2024-06-27 |
| Gemma 3 27B IT | 13 | 6/7 | 265 | 2025-03-10 |
| Gemma 2 2B IT | 4 | 5/7 | 271 | 2024-07-31 |
| Gemma 3 4B | 6 | 6/7 | 286 | 2025-03-10 |
| BERT Base Multilingual Cased | 1 | 1/7 | — | 2022-03-02 |
| DataGemma Rig 27B IT | 1 | 3/7 | — | 2024-09-12 |
| Diffusiongemma 26B A4B | 1 | 1/7 | — | 2026-06-10 |
| Flan T5 XXL | 1 | 3/7 | — | 2022-10-21 |
| Gemini 2.0 Pro | 1 | 1/7 | — | 2025-02-05 |
| Gemini 2.0 Pro Exp | 2 | 1/7 | — | 2025-02-05 |
| Gemini 3.1 Flash | 1 | 1/7 | — | — |
| Gemini 3.8 Flash Cyber | 1 | 1/7 | — | 2026-09-02 |
| Gemma 1.1 2B IT | 1 | 3/7 | — | 2024-04-05 |
| Gemma 1.1 7B IT | 2 | 4/7 | — | 2024-04-05 |
| Gemma 3 1B | 1 | 1/7 | — | 2025-03-13 |
| Gemma 3 270M | 1 | 1/7 | — | 2025-08-05 |
| Gemma 3 4B PT | 1 | 1/7 | — | 2025-02-20 |
| Gemma 3n E2B | 1 | 1/7 | — | 2025-06-26 |
| Gemma 3n E4B IT | 2 | 2/7 | — | 2025-06-03 |
| Gemma 4 12B | 2 | 4/7 | — | 2026-06-03 |
| Gemma 4 E2B | 2 | 4/7 | — | 2026-04-02 |
| Palm 2 | 2 | 4/7 | — | 2023-05-10 |
| Shieldgemma 27B | 1 | 1/7 | — | 2024-07-16 |
| T5 11B | 1 | 1/7 | — | 2022-03-02 |
| Ul2 | 1 | 1/7 | — | 2022-06-16 |
Evaluations covering Google models (115)
AA-Omniscience · AbstentionBench · Adversarial Humanities Benchmark (AHB) — Table 5 · Adversarial Poetry — AILuminate Baseline and Poetry ASR · Adversarial Robustness · Agent-SafetyBench · AgentAbstain · AgentDojo · AgentDrive Safety Compliance · AILuminate General Purpose AI Chat · AIMS Safety-Classifier Competence · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · Arena Factuality — Search Arena (factuality-only weighting) · Arena Factuality — Text Arena (factuality-only weighting) · AuAu Authoritarian Response Audit · BioSecBench-Refusal · BioTIER · BrokenMath · BullshitBench v2 · CAIS Risk Index · CheatBench direct cheating propensity · ChineseSafe · Cisco AI Defense Rolling Single-Turn Leaderboard · Claude Fable 5.1 card — Gray Swan indirect prompt injection k=15 · COMPL-AI AI-Identity Disclosure · COMPL-AI LLM RuLES Multi-Turn Rule Following · COMPL-AI TensorTrust Goal-Hijacking Resistance · Concordia AI Risk Monitor · Confabulations · Constitutional Following — Anthropic Constitution · Constitutional Following — OpenAI Model Spec · DecodingTrust · DelusionEval · DSPSafeBench · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · Every Model Cheats — Cybench Cheat Propensity · FinEval Financial Security Knowledge · FlagEval Safety and Values · FORTRESS · Google Gemini 2.5 Flash Model Card · Google Gemini 2.5 Flash-Lite Model Card · Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15 · Governance Decay under Passive Context Compaction · Gray Swan indirect prompt injection (15 attempts) · HalluVerse-M3 Hallucination Recognition · HarmBench · HarmVideoBench · HELM Classic RealToxicityPrompts · HELM Safety · HUMAINE Trust, Ethics and Safety · Human Pathogen Capabilities Test (HPCT) — overall refusal · Humanity's Last Exam RMS calibration error (Scale Labs) · IndoBias-Pairs — parity-aware culturally grounded bias · Inkling-Small model card — FORTRESS · Inkling-Small model card — StrongREJECT · JuICE Cultural-Error Span Detection · KIDBench Implicit Child Cue · kindbench v0.1.0 psychological safety ranking · Large-scale Moral Machine experiment on LLMs · LiveSecBench · LLM Ethics Benchmark · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · Microsoft Phi Safety Panels · MORU · MT-JailBench CrescendoX · MuPPET Contextual Privacy · NESSiE Necessary Safety Benchmark · ODCV-Bench · Olam Social Poker — Social Lie Rate · Opposite-Narrator Sycophancy · OR-Bench · PacifAIst · PandaBench JBB direct-request panel · Pander Score · PHARE · Pokee-Isaac model card — DTAP · PropensityBench · RealityTest — Text AI-Identity Disclosure · RefusalBench · Reward Hacking Benchmark · S-Eval · SafetyBench · SALAD-Bench · Shell · SimpleQA Verified · SM-Bench · Social Welfare Function Benchmark · SORRY-Bench · SOSBench · SpeciesismBench · SpeciEval · StereoTales Harmful Associations · SYCON Bench · TAC · ThaiSafetyBench · The Dictatorship Eval · ToolPrivacyBench · TrustLLM contemporary collapsed application · UAVBench safety-critical decision recognition · Vals AI Cheating Audit · Vectara HHEM Factual Consistency · VETO Misfired Alignment · Vigil Mental Health Safety · WildClawBench Safety & Alignment (OpenClaw harness)
