← Models

Model profile

GPT 5.4 Nano

OpenAIdeveloper
2026-03-17release date
#198 / 309overall rank
14eval lineages

Evidence summary

GPT 5.4 Nano has an estimated overall rank of #198; its 90% source-sensitivity interval is #68–#258. Its behavior-only rank is #216; company governance moves the combined estimate to #198. Published evidence spans 14 evals and 5 of 7 behavior components. Its strongest relative result is Vectara HHEM Factual Consistency (factual_consistency_rate, #2 of 94); its weakest is CAIS Risk Index (political_manipulation, #33 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#61 / 3280.5113Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#8 / 1121482.0Source ↗official
BullshitBench v2clear_pushback_rate#81 / 1060.1233Source ↗official
CAIS Risk Indexagent_red_teaming#25 / 4481.7Source ↗official
CAIS Risk Indexbioweapons_assistance#32 / 4978.3Source ↗official
CAIS Risk Indexhle_overconfidence#17 / 5047.1Source ↗official
CAIS Risk Indexmachiavelli#24 / 4688.5Source ↗official
CAIS Risk Indexmask#30 / 5223.7Source ↗official
CAIS Risk Indexpolitical_manipulation#33 / 3363.3Source ↗official
CAIS Risk Indextextquests_harm#10 / 4915.2Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#24 / 1049.521Source ↗official
MACHIAVELLIdeception_relative_random_pct#24 / 5088.5Source ↗official
SM-Benchadversarial#79 / 7968.29Source ↗official
SM-Benchambiguous_interpretation#36 / 7986.31Source ↗official
SM-Benchanti_hallucination#76 / 7968.06Source ↗official
SM-Bencheq_boundaries#52 / 7958.71Source ↗official
SM-Benchoverfit#79 / 794.37Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#2 / 9496.9Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#25 / 2518.9Source ↗official