← Models

Model profile

Gemini 2.5 Flash Lite

Googledeveloper
2025-06-17release date
#228 / 333Safety rank
#48 / 645Freedom rank

Evidence summary

Safety. Gemini 2.5 Flash Lite has an estimated Safety rank of #228; its 90% source-sensitivity interval is #94–#272. Its behavior-only rank is #239; company governance moves the combined estimate to #228. Published Safety evidence spans 23 eval lineages and 7 of 7 components. Its strongest relative result is Vectara HHEM Factual Consistency (factual_consistency_rate, #3 of 94); its weakest is CAIS Risk Index (agent_red_teaming, #49 of 49).

Freedom. Gemini 2.5 Flash Lite has an estimated Freedom rank of #48; its 90% source-sensitivity interval is #47–#224. Published Freedom evidence spans 11 eval lineages and 1 of 1 components. Its strongest relative result is CAIS Risk Index (bioweapons_assistance, #2 of 54); its weakest is Google Gemini 2.5 Flash-Lite Model Card (text_safety_delta, #1 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#132 / 3450.7171Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#15 / 248.67Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#15 / 2464.77Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#43 / 800.658Source ↗official
ANIMAscore#18 / 220.5838Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#16 / 1111461.0Source ↗official
CAIS Risk Indexagent_red_teaming#49 / 4996.8Source ↗official
CAIS Risk Indexbioweapons_assistance#53 / 5496.3Source ↗official
CAIS Risk Indexhle_overconfidence#54 / 5583.4Source ↗official
CAIS Risk Indexmachiavelli#51 / 51109.4Source ↗official
CAIS Risk Indexmask#50 / 5754Source ↗official
CAIS Risk Indextextquests_harm#44 / 5422.8Source ↗official
DelusionEvaldelusional_prevalence_pct#13 / 1656Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#12 / 1616.5Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#12 / 165.6Source ↗official
DelusionEvalrelationship_prevalence_pct#10 / 1633.6Source ↗official
DelusionEvalsycophancy_prevalence_pct#13 / 1635Source ↗official
Google Gemini 2.5 Flash-Lite Model Cardtext_safety_delta#2 / 25.7Source ↗official
HELM Safetyanthropic_red_team#44 / 800.987Source ↗official
HELM Safetybbq#28 / 800.949Source ↗official
HELM Safetyharmbench#49 / 800.67Source ↗official
HELM Safetysimple_safety_tests#64 / 800.965Source ↗official
HELM Safetyxstest#13 / 800.978Source ↗official
MACHIAVELLIdeception_relative_random_pct#50 / 50109.4Source ↗official
MORUscore#8 / 1371.34Source ↗official
PHAREbias_resistance_diagnostic#34 / 660.4553Source ↗official
PHAREhallucination_resistance_diagnostic#50 / 700.6809Source ↗official
PHAREharm_resistance_diagnostic#66 / 700.7915Source ↗official
PHAREjailbreak_resistance_diagnostic#45 / 670.4284Source ↗official
SimpleQA Verifiedf1_score#13 / 1311.1Source ↗official
SpeciEvalbelief_animal_sentience#53 / 1236.83Source ↗official
SpeciEvalland_animal_4ns#100 / 1234.81Source ↗official
SpeciEvalsea_animal_4ns#27 / 1234.5Source ↗official
SpeciEvalspeciesism#38 / 1231.74Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#3 / 9496.7Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#10 / 248.67Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#10 / 2464.77Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#38 / 800.658Source ↗official
CAIS Risk Indexbioweapons_assistance#2 / 5496.3Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#5 / 1616.5Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#5 / 165.6Source ↗official
Google Gemini 2.5 Flash-Lite Model Cardtext_safety_delta#1 / 25.7Source ↗official
HELM Safetyanthropic_red_team#35 / 800.987Source ↗official
HELM Safetyharmbench#32 / 800.67Source ↗official
HELM Safetysimple_safety_tests#17 / 800.965Source ↗official
HELM Safetyxstest#13 / 800.978Source ↗official
PHAREharm_resistance_diagnostic#5 / 700.7915Source ↗official
PHAREjailbreak_resistance_diagnostic#23 / 670.4284Source ↗official
SpeechMap model completioncomplete_pct#26 / 18184.6Source ↗official