← Models

Model profile

Gemini 2.0 Flash

Googledeveloper
2024-12-11release date
#187 / 312overall rank
13eval lineages

Evidence summary

Gemini 2.0 Flash has an estimated overall rank of #187; its 90% source-sensitivity interval is #75–#246. Its behavior-only rank is #199; company governance moves the combined estimate to #187. Published evidence spans 13 evals and 6 of 7 behavior components. Its strongest relative result is PandaBench JBB direct-request panel (safety_rate, #1 of 46); its weakest is PropensityBench (score, #13 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#210 / 3300.8624Source ↗official
AgentDojotargeted_attack_success_rate#9 / 150.1892Source ↗official
AgentDojoutility_under_attack#7 / 150.3982Source ↗official
BullshitBench v2clear_pushback_rate#74 / 1060.15Source ↗official
Confabulationsconfabulation_rate#23 / 5219.55Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#150 / 24113.7Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#35 / 24193.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#131 / 24163.89Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#169 / 23993.64Source ↗official
Google Gemini 2.5 Flash Model Cardtext_safety_delta#1 / 20Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#27 / 5427.08Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#29 / 3982Source ↗official
MASKlying_probability_pct#39 / 5349.93Source ↗official
PandaBench JBB direct-request panelsafety_rate#1 / 461Source ↗official
PHAREbias_resistance_diagnostic#18 / 660.5351Source ↗official
PHAREhallucination_resistance_diagnostic#42 / 700.7159Source ↗official
PHAREharm_resistance_diagnostic#31 / 700.943Source ↗official
PHAREjailbreak_resistance_diagnostic#56 / 670.3735Source ↗official
PropensityBenchscore#13 / 1477.8Source ↗official
SpeciEvalbelief_animal_sentience#88 / 1136.5Source ↗official
SpeciEvalland_animal_4ns#27 / 1134.29Source ↗official
SpeciEvalsea_animal_4ns#66 / 1134.79Source ↗official
SpeciEvalspeciesism#80 / 1132.33Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism67.1
Self-direction48.9
Care / Harm45.1
Fairness / Cheating42.5
Ethical88.9

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility65
Warning-phrase rate5
Soft-censorship rate0
API-error rate0