← Models

Model profile

Gemini 1.5 Flash

Googledeveloper
2024-05-14release date
#79 / 305overall rank
17eval lineages

Evidence summary

Gemini 1.5 Flash has an estimated overall rank of #79; its 90% source-sensitivity interval is #38–#179. Its behavior-only rank is #87; company governance moves the combined estimate to #79. Published evidence spans 17 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (anthropic_red_team, #2 of 80); its weakest is HELM Safety (xstest, #66 of 80).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Robustnessscore#4 / 814Source ↗official
Agent-SafetyBenchcompromise_availability#9 / 1630Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#4 / 1648.4Source ↗official
Agent-SafetyBenchleak_sensitive_information#5 / 1639.2Source ↗official
Agent-SafetyBenchphysical_harm#6 / 1638.8Source ↗official
Agent-SafetyBenchproduce_unsafe_information#11 / 1682.4Source ↗official
Agent-SafetyBenchproperty_loss#7 / 1641.6Source ↗official
Agent-SafetyBenchspread_unsafe_information#4 / 1620.8Source ↗official
Agent-SafetyBenchviolate_law_ethics#6 / 1632Source ↗official
AgentDojotargeted_attack_success_rate#4 / 150.0787Source ↗official
AgentDojoutility_under_attack#11 / 150.333Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#31 / 800.7325Source ↗official
AnimalHarmBenchscore#3 / 100.05Source ↗official
COMPL-AI AI-Identity Disclosurescore#1 / 141Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#5 / 140.5169Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#6 / 130.4923Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#183 / 24112.14Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#147 / 24186.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#154 / 24155.56Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#197 / 23990.59Source ↗official
FinEval Financial Security Knowledgefinancial_security_accuracy_pct#6 / 1977.5Source ↗official
FORTRESSaverage_risk_score#35 / 4950.61Source ↗official
FORTRESSover_refusal_score#22 / 464.45Source ↗official
HELM Safetyanthropic_red_team#2 / 800.999Source ↗official
HELM Safetybbq#32 / 800.947Source ↗official
HELM Safetyharmbench#35 / 800.8Source ↗official
HELM Safetysimple_safety_tests#58 / 800.97Source ↗official
HELM Safetyxstest#66 / 800.921Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#23 / 391.116Source ↗official
OR-Benchover_refusal_rate#17 / 2584.3Source ↗official
OR-Benchtoxic_acceptance_rate#8 / 251.2Source ↗official
SORRY-Benchavg#5 / 510.08Source ↗official