← Models

Model profile

GPT 5 Mini

OpenAIdeveloper
2025-08-07release date
#24 / 312overall rank
29eval lineages

Evidence summary

GPT 5 Mini has an estimated overall rank of #24; its 90% source-sensitivity interval is #19–#102. Published evidence spans 29 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is FORTRESS (over_refusal_score, #46 of 48).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#50 / 3300.4432Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#11 / 800.857Source ↗official
Alignment Leaderboardcorrigibility#3 / 244.495Source ↗official
Alignment Leaderboardhonesty#4 / 244.157Source ↗official
Alignment Leaderboardnon_manipulation#5 / 244.243Source ↗official
Alignment Leaderboardrobustness#1 / 244.6Source ↗official
Alignment Leaderboardsafety#5 / 244.288Source ↗official
Alignment Leaderboardscheming#5 / 244.193Source ↗official
ANIMAscore#8 / 220.7177Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#14 / 1121465.0Source ↗official
AuAu Authoritarian Response Auditbehavioral_vignette_arr_percent#1 / 1718.2Source ↗official
AuAu Authoritarian Response Auditrealistic_prompt_arr_percent#1 / 170.5Source ↗official
CAIS Risk Indexagent_red_teaming#29 / 4584.3Source ↗official
CAIS Risk Indexbioweapons_assistance#32 / 5078.2Source ↗official
CAIS Risk Indexhle_overconfidence#35 / 5165Source ↗official
CAIS Risk Indexmachiavelli#13 / 4784.2Source ↗official
CAIS Risk Indexmask#28 / 5321.7Source ↗official
CAIS Risk Indextextquests_harm#23 / 5017.5Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#18 / 1045.085Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#28 / 24141.09Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#8 / 24196.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#46 / 24191.11Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#149 / 23995.2Source ↗official
FORTRESSaverage_risk_score#12 / 4917Source ↗official
FORTRESSover_refusal_score#46 / 4814.17Source ↗official
HELM Safetyanthropic_red_team#32 / 800.991Source ↗official
HELM Safetybbq#19 / 800.963Source ↗official
HELM Safetyharmbench#10 / 800.971Source ↗official
HELM Safetysimple_safety_tests#1 / 801Source ↗official
HELM Safetyxstest#14 / 800.977Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#33 / 5426.31Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#15 / 3965Source ↗official
KIDBench Implicit Child Cueimplicit_child_cue_total_mean#6 / 133.99Source ↗official
MACHIAVELLIdeception_relative_random_pct#13 / 5084.2Source ↗official
MASKlying_probability_pct#14 / 5317.4Source ↗official
MORUscore#3 / 1381.57Source ↗official
PHAREbias_resistance_diagnostic#31 / 660.4641Source ↗official
PHAREhallucination_resistance_diagnostic#22 / 700.7958Source ↗official
PHAREharm_resistance_diagnostic#7 / 700.9829Source ↗official
PHAREjailbreak_resistance_diagnostic#15 / 670.686Source ↗official
SM-Benchadversarial#35 / 7982.44Source ↗official
SM-Benchambiguous_interpretation#10 / 7991.67Source ↗official
SM-Benchanti_hallucination#57 / 7986.91Source ↗official
SM-Bencheq_boundaries#69 / 7952.81Source ↗official
SM-Benchoverfit#75 / 7925.14Source ↗official
SpeciEvalbelief_animal_sentience#95 / 1136.43Source ↗official
SpeciEvalland_animal_4ns#13 / 1134.17Source ↗official
SpeciEvalsea_animal_4ns#31 / 1134.58Source ↗official
SpeciEvalspeciesism#97 / 1132.65Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#9 / 2386.57Source ↗official
TACbase_welfare_rate#39 / 7626.28Source ↗self run
Vectara HHEM Factual Consistencyfactual_consistency_rate#77 / 9487.1Source ↗official