← Models

Model profile

Gemini 1.0 Pro

Googledeveloper
2023-12-06release date
#57 / 309overall rank
8eval lineages

Evidence summary

Gemini 1.0 Pro has an estimated overall rank of #57; its 90% source-sensitivity interval is #18–#194. Its behavior-only rank is #64; company governance moves the combined estimate to #57. Published evidence spans 8 evals and 6 of 7 behavior components. Its strongest relative result is OR-Bench (over_refusal_rate, #2 of 25); its weakest is OR-Bench (toxic_acceptance_rate, #22 of 25).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#53 / 800.582Source ↗official
DecodingTrustmachine_ethics#1 / 893.74Source ↗official
DecodingTruststereotype_bias#2 / 898.33Source ↗official
DecodingTrusttoxicity#3 / 877.53Source ↗official
HarmBenchdr#11 / 2818Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#12 / 390.8746Source ↗official
OR-Benchover_refusal_rate#2 / 259.7Source ↗official
OR-Benchtoxic_acceptance_rate#22 / 2521.3Source ↗official
S-Evalbase_en_overall#19 / 2241.9Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#16 / 3313.36Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#12 / 3320.85Source ↗official
SALAD-Benchattack_enhanced_malicious_use#16 / 3312.23Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#15 / 3317.93Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#8 / 3327.13Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#16 / 3312.99Source ↗official
SALAD-Benchbase_human_autonomy_integrity#22 / 3391.09Source ↗official
SALAD-Benchbase_information_safety_harms#24 / 3391.4Source ↗official
SALAD-Benchbase_malicious_use#23 / 3387.27Source ↗official
SALAD-Benchbase_misinformation_harms#16 / 3393.75Source ↗official
SALAD-Benchbase_representation_toxicity#22 / 3387.37Source ↗official
SALAD-Benchbase_socioeconomic_harms#24 / 3382.26Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#11 / 3345.28Source ↗official
SALAD-Benchmcq_information_safety_harms#11 / 3343.33Source ↗official
SALAD-Benchmcq_malicious_use#11 / 3346.6Source ↗official
SALAD-Benchmcq_misinformation_harms#11 / 3344.76Source ↗official
SALAD-Benchmcq_representation_toxicity#15 / 3339.9Source ↗official
SALAD-Benchmcq_socioeconomic_harms#11 / 3344.44Source ↗official
SORRY-Benchavg#29 / 510.33Source ↗official