← Models

Model profile

GPT 5.4 Mini

OpenAIdeveloper
2026-03-17release date
#168 / 267overall rank
14eval lineages

Evidence summary

GPT 5.4 Mini has an estimated overall rank of #168; its 90% source-sensitivity interval is #54–#218. Its behavior-only rank is #178; company governance moves the combined estimate to #168. Published evidence spans 14 evals and 6 of 7 behavior components. Its strongest relative result is DystopiaBench (petrov_score, #9 of 50); its weakest is SM-Bench (overfit, #71 of 73).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#244 / 3110.8916↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#49 / 1050.3167↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#21 / 4374.3↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#12 / 4860↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#23 / 4955.2↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#39 / 4595.6↓ lowerSource ↗official
CAIS Risk Indexmask#21 / 5115.9↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#30 / 3261.1↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#12 / 4815.4↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#26 / 1059.571↓ lowerSource ↗official
DystopiaBenchbasaglia_score#14 / 5049.13↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#12 / 5038.87↓ lowerSource ↗official
DystopiaBenchhuxley_score#11 / 5045.1↓ lowerSource ↗official
DystopiaBenchlaguardia_score#11 / 5052.8↓ lowerSource ↗official
DystopiaBenchorwell_score#12 / 5041.8↓ lowerSource ↗official
DystopiaBenchpetrov_score#9 / 5039.6↓ lowerSource ↗official
MACHIAVELLIdeception_relative_random_pct#41 / 5095.6↓ lowerSource ↗official
RefusalBenchyouden_j#16 / 190.06411↑ higherSource ↗official
SM-Benchadversarial#60 / 7376.83↑ higherSource ↗official
SM-Benchambiguous_interpretation#38 / 7384.08↑ higherSource ↗official
SM-Benchanti_hallucination#63 / 7380.23↑ higherSource ↗official
SM-Bencheq_boundaries#28 / 7367.28↑ higherSource ↗official
SM-Benchoverfit#71 / 7315.3↑ higherSource ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#17 / 259.9↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience5.5
Honesty-humility5.3
Extraversion6.1
Agreeableness6.5
Conscientiousness7.3

Agent-ValueBench Schwartz Basic Values (PVQ40)