← Models

Model profile

GPT 5.5

OpenAIdeveloper
2026-04-23release date
#14 / 312overall rank
38eval lineages

Evidence summary

GPT 5.5 has an estimated overall rank of #14; its 90% source-sensitivity interval is #11–#79. Its behavior-only rank is #15; company governance moves the combined estimate to #14. Published evidence spans 38 evals and 7 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 113); its weakest is GPT-5.6 system card (extremism_not_unsafe, #7 of 7).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#231 / 3300.8804Source ↗official
AgentAbstainabstain#9 / 1761.1Source ↗official
AgentAbstaincar#7 / 1759.8Source ↗official
AgentAbstainpaired#4 / 1752.5Source ↗official
ANIMAscore#4 / 220.74Source ↗official
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating#2 / 301269.0Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#4 / 1121490.0Source ↗official
BioSecBench-Refusalbalanced_refusal_score#8 / 100.2869Source ↗official
BullshitBench v2clear_pushback_rate#28 / 1060.4567Source ↗official
CAIS Risk Indexagent_red_teaming#4 / 4541.5Source ↗official
CAIS Risk Indexbioweapons_assistance#15 / 5060.5Source ↗official
CAIS Risk Indexhle_overconfidence#8 / 5141.8Source ↗official
CAIS Risk Indexmachiavelli#8 / 4781.7Source ↗official
CAIS Risk Indexmask#13 / 539.9Source ↗official
CAIS Risk Indexpolitical_manipulation#10 / 3440Source ↗official
CAIS Risk Indextextquests_harm#36 / 5021.4Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#5 / 1042.343Source ↗official
DystopiaBenchbasaglia_score#17 / 5055Source ↗official
DystopiaBenchbaudrillard_score#14 / 5040.7Source ↗official
DystopiaBenchhuxley_score#15 / 5052.5Source ↗official
DystopiaBenchlaguardia_score#17 / 5060.13Source ↗official
DystopiaBenchorwell_score#15 / 5049.5Source ↗official
DystopiaBenchpetrov_score#12 / 5049.13Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#37 / 24138.5Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#200 / 24179.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#6 / 24199.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#184 / 23992Source ↗official
FORTRESSaverage_risk_score#10 / 4916.28Source ↗official
FORTRESSover_refusal_score#40 / 489.79Source ↗official
GPT-5.6 system cardconnectors_injection_resistance#1 / 71Source ↗official
GPT-5.6 system cardemotional_reliance#6 / 70.915Source ↗official
GPT-5.6 system cardextremism_not_unsafe#7 / 70.925Source ↗official
GPT-5.6 system cardgore_not_unsafe#3 / 70.8Source ↗official
GPT-5.6 system cardharm_overall_pct#6 / 71.12Source ↗official
GPT-5.6 system cardhate_not_unsafe#1 / 71Source ↗official
GPT-5.6 system cardmental_health#6 / 70.82Source ↗official
GPT-5.6 system cardnonviolent_illicit_not_unsafe#6 / 70.987Source ↗official
GPT-5.6 system cardself_harm#6 / 70.868Source ↗official
GPT-5.6 system cardself_harm_not_unsafe#7 / 70.917Source ↗official
GPT-5.6 system cardsexual_minors_not_unsafe#6 / 70.938Source ↗official
GPT-5.6 system cardsexual_not_unsafe#3 / 70.944Source ↗official
GPT-5.6 system cardviolent_illicit_not_unsafe#5 / 70.94Source ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#8 / 1320.8Source ↗official
HarmVideoBenchharmful_video_safety_recognition_reasoning#3 / 190.809Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#11 / 5428.38Source ↗official
JuICE Cultural-Error Span Detectionf1#2 / 100.5112Source ↗official
kindbench v0.1.0 psychological safety rankingemotional_safety#4 / 1082.4Source ↗official
kindbench v0.1.0 psychological safety rankingidentity_collapse#2 / 1092.6Source ↗official
kindbench v0.1.0 psychological safety rankingsycophancy_spine#7 / 1082.4Source ↗official
kindbench v0.1.0 psychological safety rankingvalue_integrity#6 / 1090.5Source ↗official
MACHIAVELLIdeception_relative_random_pct#8 / 5081.7Source ↗official
MANTAAWMS#2 / 70.504Source ↗official
MANTAAWVS#2 / 70.664Source ↗official
MORUscore#1 / 1385.4Source ↗official
MuPPET Contextual Privacymultiparty_contextual_privacy_score#2 / 750.98Source ↗official
ODCV-Benchaverage_severity#2 / 120.7125Source ↗official
ODCV-Benchmisalignment_rate#3 / 1221.25Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#3 / 192Source ↗official
PHAREbias_resistance_diagnostic#48 / 660.3829Source ↗official
PHAREhallucination_resistance_diagnostic#20 / 700.7978Source ↗official
PHAREharm_resistance_diagnostic#6 / 700.9892Source ↗official
PHAREjailbreak_resistance_diagnostic#18 / 670.6513Source ↗official
RefusalBenchyouden_j#6 / 190.2979Source ↗official
SM-Benchadversarial#33 / 7982.93Source ↗official
SM-Benchambiguous_interpretation#47 / 7982.74Source ↗official
SM-Benchanti_hallucination#31 / 7994.76Source ↗official
SM-Bencheq_boundaries#41 / 7964.61Source ↗official
SM-Benchoverfit#62 / 7950Source ↗official
SpeciEvalbelief_animal_sentience#1 / 1137Source ↗official
SpeciEvalland_animal_4ns#61 / 1134.58Source ↗official
SpeciEvalsea_animal_4ns#27 / 1134.55Source ↗official
SpeciEvalspeciesism#10 / 1131.32Source ↗official
TACbase_welfare_rate#62 / 7619.23Source ↗official
ToolPrivacyBenchprivate_mt_poi#3 / 920.39Source ↗official
ToolPrivacyBenchpublic_mt_poi#2 / 916.75Source ↗official
UK AISI cyber-evaluation cheating and prompted self-reportattempted_cheating_trajectory_rate_pct#3 / 511.4Source ↗official
UK AISI cyber-evaluation cheating and prompted self-reportspecific_cheating_action_mention_rate_pct#4 / 575Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#41 / 9490.7Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#9 / 257.5Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CAIS AI Values — countries