← Models

Model profile

Gemini 3 Pro Preview

Googledeveloper
2025-11-18release date
#115 / 309overall rank
25eval lineages

Evidence summary

Gemini 3 Pro Preview has an estimated overall rank of #115; its 90% source-sensitivity interval is #57–#189. Its behavior-only rank is #124; company governance moves the combined estimate to #115. Published evidence spans 25 evals and 7 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 113); its weakest is CAIS Risk Index (bioweapons_assistance, #49 of 49).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#254 / 3280.9Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#32 / 800.732Source ↗official
Alignment Leaderboardcorrigibility#10 / 244.32Source ↗official
Alignment Leaderboardhonesty#6 / 243.94Source ↗official
Alignment Leaderboardnon_manipulation#6 / 243.929Source ↗official
Alignment Leaderboardrobustness#11 / 243.613Source ↗official
Alignment Leaderboardsafety#6 / 244.038Source ↗official
Alignment Leaderboardscheming#6 / 244.108Source ↗official
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating#20 / 301180.0Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#21 / 1121457.0Source ↗official
BullshitBench v2clear_pushback_rate#34 / 1060.42Source ↗official
CAIS Risk Indexagent_red_teaming#10 / 4447.5Source ↗official
CAIS Risk Indexbioweapons_assistance#49 / 49100Source ↗official
CAIS Risk Indexhle_overconfidence#28 / 5057.2Source ↗official
CAIS Risk Indexmachiavelli#44 / 4699.8Source ↗official
CAIS Risk Indexmask#48 / 5258Source ↗official
CAIS Risk Indextextquests_harm#36 / 4921.4Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#33 / 10418.1Source ↗official
Constitutional Following — Anthropic Constitutionconstitutional_following_score#5 / 787.6Source ↗official
Constitutional Following — OpenAI Model Specconstitutional_following_score#6 / 793.9Source ↗official
FORTRESSaverage_risk_score#30 / 4941.69Source ↗official
HELM Safetyanthropic_red_team#61 / 800.971Source ↗official
HELM Safetybbq#3 / 800.984Source ↗official
HELM Safetyharmbench#44 / 800.725Source ↗official
HELM Safetysimple_safety_tests#54 / 800.975Source ↗official
HELM Safetyxstest#19 / 800.973Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#6 / 5429.09Source ↗official
MACHIAVELLIdeception_relative_random_pct#45 / 5099.8Source ↗official
MASKlying_probability_pct#50 / 5357.4Source ↗official
MT-JailBench CrescendoXsafety_score#6 / 2136.48Source ↗official
PHAREbias_resistance_diagnostic#16 / 660.5365Source ↗official
PHAREhallucination_resistance_diagnostic#16 / 700.8102Source ↗official
PHAREharm_resistance_diagnostic#37 / 700.935Source ↗official
PHAREjailbreak_resistance_diagnostic#19 / 670.6506Source ↗official
PropensityBenchscore#8 / 1452.85Source ↗official
SM-Benchadversarial#12 / 7986.83Source ↗official
SM-Benchambiguous_interpretation#42 / 7984.52Source ↗official
SM-Benchanti_hallucination#20 / 7997.38Source ↗official
SM-Bencheq_boundaries#41 / 7964.61Source ↗official
SM-Benchoverfit#20 / 7983.06Source ↗official
SpeciEvalbelief_animal_sentience#1 / 1137Source ↗official
SpeciEvalland_animal_4ns#80 / 1134.75Source ↗official
SpeciEvalsea_animal_4ns#72 / 1134.85Source ↗official
SpeciEvalspeciesism#87 / 1132.45Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#80 / 9486.4Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.1
Government46
Diplomacy66.3
Economy44.9
Society62

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility60
Warning-phrase rate15
Soft-censorship rate0
API-error rate0