← Models

Model profile

Gemini 2.5 Pro Exp

Googledeveloper
2025-03-25release date
#38 / 333Safety rank
#140 / 645Freedom rank

Evidence summary

Safety. Gemini 2.5 Pro Exp has an estimated Safety rank of #38; its 90% source-sensitivity interval is #4–#195. Its behavior-only rank is #39; company governance moves the combined estimate to #38. Published Safety evidence spans 4 eval lineages and 3 of 7 components. Its strongest relative result is Confabulations (confabulation_rate, #3 of 52); its weakest is FORTRESS (average_risk_score, #40 of 50).

Freedom. Gemini 2.5 Pro Exp has an estimated Freedom rank of #140; its 90% source-sensitivity interval is #86–#430. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is FORTRESS (over_refusal_score, #6 of 49); its weakest is FORTRESS (average_risk_score, #11 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
CAIS Risk Indexhle_overconfidence#42 / 5571Source ↗official
CAIS Risk Indexmask#44 / 5746.9Source ↗official
Confabulationsconfabulation_rate#3 / 523.96Source ↗official
FORTRESSaverage_risk_score#40 / 5054.89Source ↗official
FORTRESSover_refusal_score#6 / 491.38Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#22 / 4271Source ↗official
MASKlying_probability_pct#33 / 5344.07Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
FORTRESSaverage_risk_score#11 / 5054.89Source ↗official
FORTRESSover_refusal_score#6 / 491.38Source ↗official