← Models

Model profile

Gemini 1.5 Pro

Googledeveloper
2024-02-15release date
#70 / 267overall rank
20eval lineages

Evidence summary

Gemini 1.5 Pro has an estimated overall rank of #70; its 90% source-sensitivity interval is #29–#160. Its behavior-only rank is #77; company governance moves the combined estimate to #70. Published evidence spans 20 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (anthropic_red_team, #2 of 80); its weakest is HELM Safety (xstest, #69 of 80).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AbstentionBenchanswer_unknown_f1#2 / 200.9132↑ higherSource ↗official
AbstentionBenchfalse_premise_f1#3 / 200.7501↑ higherSource ↗official
AbstentionBenchstale_f1#17 / 200.5854↑ higherSource ↗official
AbstentionBenchsubjective_f1#6 / 200.7585↑ higherSource ↗official
AbstentionBenchunderspecified_context_f1#4 / 200.6963↑ higherSource ↗official
AbstentionBenchunderspecified_intent_f1#9 / 200.7403↑ higherSource ↗official
Adversarial Robustnessscore#1 / 88↓ lowerSource ↗official
Agent-SafetyBenchcompromise_availability#8 / 1630.8↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#5 / 1642↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#10 / 1630↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#10 / 1628.8↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#10 / 1684.8↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#9 / 1637.6↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#5 / 1618.8↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#9 / 1626.8↑ higherSource ↗official
AgentDojotargeted_attack_success_rate#10 / 150.2282↓ lowerSource ↗official
AgentDojoutility_under_attack#8 / 150.38↑ higherSource ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#27 / 3234.7↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#22 / 3221.6↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#21 / 3211.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#19 / 3226.7↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#17 / 3213.3↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#18 / 3217.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#19 / 3214.6↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#18 / 3216.5↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#16 / 3223.4↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#16 / 3215.6↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#18 / 3213.6↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#19 / 3220.6↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#26 / 800.7505↑ higherSource ↗official
AnimalHarmBenchscore#2 / 100.066↑ higherSource ↗official
Confabulationsconfabulation_rate#18 / 5216.83↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#148 / 26014.73↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#94 / 26090.17↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#138 / 26068.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#159 / 25895.41↑ higherSource ↗official
FORTRESSaverage_risk_score#38 / 4953.86↓ lowerSource ↗official
FORTRESSover_refusal_score#19 / 463.17↓ lowerSource ↗official
HELM Safetyanthropic_red_team#2 / 800.999↑ higherSource ↗official
HELM Safetybbq#33 / 800.945↑ higherSource ↗official
HELM Safetyharmbench#36 / 800.799↑ higherSource ↗official
HELM Safetysimple_safety_tests#54 / 800.975↑ higherSource ↗official
HELM Safetyxstest#69 / 800.904↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#18 / 391.009↓ lowerSource ↗official
OR-Benchover_refusal_rate#19 / 2588↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#7 / 250.6↓ lowerSource ↗official
PHAREhallucination_resistance_diagnostic#26 / 700.7807↑ higherSource ↗official
PHAREharm_resistance_diagnostic#16 / 700.9684↑ higherSource ↗official
SORRY-Benchavg#2 / 510.07↓ lowerSource ↗official
SpeciesismBenchexplicit_speciesism_scale#2 / 72.117↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism68.8
Self-direction53.2
Care / Harm25.7
Fairness / Cheating21.4
Ethical88.5