← Models

Model profile

Gemini 1.5 Pro

Googledeveloper
2024-02-15release date
#68 / 312overall rank
22eval lineages

Evidence summary

Gemini 1.5 Pro has an estimated overall rank of #68; its 90% source-sensitivity interval is #32–#176. Its behavior-only rank is #76; company governance moves the combined estimate to #68. Published evidence spans 22 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (anthropic_red_team, #2 of 80); its weakest is Humanity's Last Exam RMS calibration error (Scale Labs) (calibrationError, #37 of 39).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AbstentionBenchanswer_unknown_f1#2 / 200.9132Source ↗official
AbstentionBenchfalse_premise_f1#3 / 200.7501Source ↗official
AbstentionBenchstale_f1#17 / 200.5854Source ↗official
AbstentionBenchsubjective_f1#6 / 200.7585Source ↗official
AbstentionBenchunderspecified_context_f1#4 / 200.6963Source ↗official
AbstentionBenchunderspecified_intent_f1#9 / 200.7403Source ↗official
Adversarial Robustnessscore#1 / 88Source ↗official
Agent-SafetyBenchcompromise_availability#8 / 1630.8Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#5 / 1642Source ↗official
Agent-SafetyBenchleak_sensitive_information#10 / 1630Source ↗official
Agent-SafetyBenchphysical_harm#10 / 1628.8Source ↗official
Agent-SafetyBenchproduce_unsafe_information#10 / 1684.8Source ↗official
Agent-SafetyBenchproperty_loss#9 / 1637.6Source ↗official
Agent-SafetyBenchspread_unsafe_information#5 / 1618.8Source ↗official
Agent-SafetyBenchviolate_law_ethics#9 / 1626.8Source ↗official
AgentDojotargeted_attack_success_rate#10 / 150.2282Source ↗official
AgentDojoutility_under_attack#8 / 150.38Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#27 / 3234.7Source ↗official
AILuminate General Purpose AI Chatdefamation#22 / 3221.6Source ↗official
AILuminate General Purpose AI Chathate#21 / 3211.2Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#19 / 3226.7Source ↗official
AILuminate General Purpose AI Chatintellectual_property#17 / 3213.3Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#18 / 3217.9Source ↗official
AILuminate General Purpose AI Chatprivacy#19 / 3214.6Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#18 / 3216.5Source ↗official
AILuminate General Purpose AI Chatsexual_content#16 / 3223.4Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#16 / 3215.6Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#18 / 3213.6Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#19 / 3220.6Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#26 / 800.7505Source ↗official
AnimalHarmBenchscore#2 / 100.066Source ↗official
Confabulationsconfabulation_rate#18 / 5216.83Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#136 / 24114.73Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#84 / 24190.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#123 / 24168.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#145 / 23995.41Source ↗official
FinEval Financial Security Knowledgefinancial_security_accuracy_pct#5 / 1977.8Source ↗official
FORTRESSaverage_risk_score#38 / 4953.86Source ↗official
FORTRESSover_refusal_score#20 / 483.17Source ↗official
HELM Safetyanthropic_red_team#2 / 800.999Source ↗official
HELM Safetybbq#33 / 800.945Source ↗official
HELM Safetyharmbench#36 / 800.799Source ↗official
HELM Safetysimple_safety_tests#54 / 800.975Source ↗official
HELM Safetyxstest#69 / 800.904Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#37 / 3988Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#18 / 391.009Source ↗official
OR-Benchover_refusal_rate#19 / 2588Source ↗official
OR-Benchtoxic_acceptance_rate#7 / 250.6Source ↗official
PHAREhallucination_resistance_diagnostic#26 / 700.7807Source ↗official
PHAREharm_resistance_diagnostic#16 / 700.9684Source ↗official
SORRY-Benchavg#2 / 510.07Source ↗official
SpeciesismBenchexplicit_speciesism_scale#2 / 72.117Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism68.8
Self-direction53.2
Care / Harm25.7
Fairness / Cheating21.4
Ethical88.5