← Models

Model profile

GPT 5 Nano

OpenAIdeveloper
2025-08-07release date
#11 / 267overall rank
20eval lineages

Evidence summary

GPT 5 Nano has an estimated overall rank of #11; its 90% source-sensitivity interval is #5–#55. Its behavior-only rank is #8; company governance moves the combined estimate to #11. Published evidence spans 20 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is CAIS Risk Index (hle_overconfidence, #44 of 49).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#48 / 3110.4885↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#6 / 800.878↑ higherSource ↗official
ANIMAscore#9 / 190.708↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#31 / 4388.5↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#27 / 4871.3↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#44 / 4980↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#1 / 4579.4↓ lowerSource ↗official
CAIS Risk Indexmask#16 / 5112.3↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#1 / 482↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#21 / 1057.527↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#36 / 26039.02↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#4 / 26098↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#45 / 26092.78↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#116 / 25896.91↑ higherSource ↗official
HELM Safetyanthropic_red_team#13 / 800.996↑ higherSource ↗official
HELM Safetybbq#8 / 800.976↑ higherSource ↗official
HELM Safetyharmbench#3 / 800.984↑ higherSource ↗official
HELM Safetysimple_safety_tests#1 / 801↑ higherSource ↗official
HELM Safetyxstest#16 / 800.976↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#1 / 5079.4↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#5 / 315.067↓ lowerSource ↗self run
PHAREbias_resistance_diagnostic#55 / 660.347↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#31 / 700.7637↑ higherSource ↗official
PHAREharm_resistance_diagnostic#9 / 700.9741↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#12 / 670.6964↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#88 / 1026.4↑ higherSource ↗official
SpeciEvalland_animal_4ns#27 / 1024.35↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#40 / 1024.68↓ lowerSource ↗official
SpeciEvalspeciesism#71 / 1022.33↓ lowerSource ↗official
TACbase_welfare_rate#10 / 6837.82↑ higherSource ↗self run