← Models

Model profile

Gemini 3 Pro Preview

Googledeveloper
2025-11-18release date
#83 / 267overall rank
22eval lineages

Evidence summary

Gemini 3 Pro Preview has an estimated overall rank of #83; its 90% source-sensitivity interval is #20–#160. Its behavior-only rank is #89; company governance moves the combined estimate to #83. Published evidence spans 22 evals and 7 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 102); its weakest is CAIS Risk Index (bioweapons_assistance, #48 of 48).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#196 / 3110.8486↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#32 / 800.732↑ higherSource ↗official
Alignment Leaderboardcorrigibility#10 / 244.32↑ higherSource ↗official
Alignment Leaderboardhonesty#6 / 243.94↑ higherSource ↗official
Alignment Leaderboardnon_manipulation#6 / 243.929↑ higherSource ↗official
Alignment Leaderboardrobustness#11 / 243.613↑ higherSource ↗official
Alignment Leaderboardsafety#6 / 244.038↑ higherSource ↗official
Alignment Leaderboardscheming#6 / 244.108↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#34 / 1050.42↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#9 / 4347.5↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#48 / 48100↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#27 / 4957.2↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#43 / 4599.8↓ lowerSource ↗official
CAIS Risk Indexmask#47 / 5158↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#36 / 4821.4↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#33 / 10518.1↓ lowerSource ↗official
FORTRESSaverage_risk_score#30 / 4941.69↓ lowerSource ↗official
HELM Safetyanthropic_red_team#61 / 800.971↑ higherSource ↗official
HELM Safetybbq#3 / 800.984↑ higherSource ↗official
HELM Safetyharmbench#44 / 800.725↑ higherSource ↗official
HELM Safetysimple_safety_tests#54 / 800.975↑ higherSource ↗official
HELM Safetyxstest#19 / 800.973↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#6 / 5429.09↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#45 / 5099.8↓ lowerSource ↗official
MASKlying_probability_pct#50 / 5357.4↓ lowerSource ↗official
PHAREbias_resistance_diagnostic#16 / 660.5365↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#16 / 700.8102↑ higherSource ↗official
PHAREharm_resistance_diagnostic#37 / 700.935↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#19 / 670.6506↑ higherSource ↗official
PropensityBenchscore#8 / 1452.85↓ lowerSource ↗official
SM-Benchadversarial#12 / 7386.83↑ higherSource ↗official
SM-Benchambiguous_interpretation#37 / 7384.52↑ higherSource ↗official
SM-Benchanti_hallucination#17 / 7397.38↑ higherSource ↗official
SM-Bencheq_boundaries#36 / 7364.61↑ higherSource ↗official
SM-Benchoverfit#18 / 7383.06↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#1 / 1027↑ higherSource ↗official
SpeciEvalland_animal_4ns#72 / 1024.75↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#63 / 1024.85↓ lowerSource ↗official
SpeciEvalspeciesism#78 / 1022.45↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.1
Government46
Diplomacy66.3
Economy44.9
Society62

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility60
Warning-phrase rate15
Soft-censorship rate0
API-error rate0