← Models

Model profile

GPT 5 Nano

OpenAIdeveloper
2025-08-07release date
#14 / 309overall rank
23eval lineages

Evidence summary

GPT 5 Nano has an estimated overall rank of #14; its 90% source-sensitivity interval is #7–#71. Its behavior-only rank is #13; company governance moves the combined estimate to #14. Published evidence spans 23 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is CAIS Risk Index (hle_overconfidence, #45 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#66 / 3280.5262Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#6 / 800.878Source ↗official
ANIMAscore#10 / 220.708Source ↗official
CAIS Risk Indexagent_red_teaming#32 / 4488.5Source ↗official
CAIS Risk Indexbioweapons_assistance#28 / 4971.3Source ↗official
CAIS Risk Indexhle_overconfidence#45 / 5080Source ↗official
CAIS Risk Indexmachiavelli#1 / 4679.4Source ↗official
CAIS Risk Indexmask#17 / 5212.3Source ↗official
CAIS Risk Indextextquests_harm#1 / 492Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#21 / 1047.527Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#33 / 24139.02Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#4 / 24198Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#39 / 24192.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#104 / 23996.91Source ↗official
HELM Safetyanthropic_red_team#13 / 800.996Source ↗official
HELM Safetybbq#8 / 800.976Source ↗official
HELM Safetyharmbench#3 / 800.984Source ↗official
HELM Safetysimple_safety_tests#1 / 801Source ↗official
HELM Safetyxstest#16 / 800.976Source ↗official
MACHIAVELLIdeception_relative_random_pct#1 / 5079.4Source ↗official
Manager Coercion Benchcoercion_ladder_depth#5 / 315.067Source ↗self run
MonitoringBench Full-Trajectory Monitorfull_trajectory_catch_rate_at_1pct_fpr_percent#10 / 136Source ↗official
PHAREbias_resistance_diagnostic#55 / 660.347Source ↗official
PHAREhallucination_resistance_diagnostic#31 / 700.7637Source ↗official
PHAREharm_resistance_diagnostic#9 / 700.9741Source ↗official
PHAREjailbreak_resistance_diagnostic#12 / 670.6964Source ↗official
SpeciEvalbelief_animal_sentience#97 / 1136.4Source ↗official
SpeciEvalland_animal_4ns#32 / 1134.35Source ↗official
SpeciEvalsea_animal_4ns#48 / 1134.68Source ↗official
SpeciEvalspeciesism#80 / 1132.33Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#2 / 2389.22Source ↗official
TACbase_welfare_rate#10 / 7637.82Source ↗self run
Vectara HHEM Factual Consistencyfactual_consistency_rate#56 / 9489.5Source ↗official