← Models

Model profile

GPT 5.4 Nano

OpenAIdeveloper
2026-03-17release date
#201 / 333Safety rank
#605 / 645Freedom rank

Evidence summary

Safety. GPT 5.4 Nano has an estimated Safety rank of #201; its 90% source-sensitivity interval is #22–#274. Its behavior-only rank is #218; company governance moves the combined estimate to #201. Published Safety evidence spans 16 eval lineages and 6 of 7 components. Its strongest relative result is Vectara HHEM Factual Consistency (factual_consistency_rate, #2 of 94); its weakest is SM-Bench (overfit, #84 of 84).

Freedom. GPT 5.4 Nano has an estimated Freedom rank of #605; its 90% source-sensitivity interval is #365–#641. Published Freedom evidence spans 6 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (adversarial, #1 of 84); its weakest is SM-Bench (overfit, #84 of 84).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#79 / 3450.5113Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#8 / 1111482.0Source ↗official
BullshitBench v2clear_pushback_rate#92 / 1170.1267Source ↗official
CAIS Risk Indexagent_red_teaming#30 / 4981.7Source ↗official
CAIS Risk Indexbioweapons_assistance#37 / 5478.3Source ↗official
CAIS Risk Indexhle_overconfidence#20 / 5547.1Source ↗official
CAIS Risk Indexmachiavelli#29 / 5188.5Source ↗official
CAIS Risk Indexmask#33 / 5723.7Source ↗official
CAIS Risk Indexpolitical_manipulation#46 / 4863.3Source ↗official
CAIS Risk Indextextquests_harm#11 / 5415.2Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#24 / 1049.521Source ↗official
DelusionEvaldelusional_prevalence_pct#4 / 1618Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#3 / 1652.2Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#1 / 160Source ↗official
DelusionEvalrelationship_prevalence_pct#4 / 1611.8Source ↗official
DelusionEvalsycophancy_prevalence_pct#3 / 1615.6Source ↗official
MACHIAVELLIdeception_relative_random_pct#24 / 5088.5Source ↗official
SM-Benchadversarial#84 / 8468.29Source ↗official
SM-Benchambiguous_interpretation#41 / 8486.31Source ↗official
SM-Benchanti_hallucination#81 / 8468.06Source ↗official
SM-Bencheq_boundaries#57 / 8458.71Source ↗official
SM-Benchoverfit#84 / 844.37Source ↗official
Vals AI Cheating Auditterminal_bench_cheating_shortcut_evidence_rate_pct#4 / 140.7491Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#2 / 9496.9Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#25 / 2518.9Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
CAIS Risk Indexbioweapons_assistance#17 / 5478.3Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#80 / 1049.521Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#14 / 1652.2Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#15 / 160Source ↗official
SM-Benchadversarial#1 / 8468.29Source ↗official
SM-Bencheq_boundaries#57 / 8458.71Source ↗official
SM-Benchoverfit#84 / 844.37Source ↗official
SpeechMap model completioncomplete_pct#167 / 18132.2Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#25 / 2518.9Source ↗official