← Models

Model profile

GPT-5.3 Chat

OpenAIdeveloper
2026-03-03release date
#9 / 309overall rank
6eval lineages

Evidence summary

GPT-5.3 Chat has an estimated overall rank of #9; its 90% source-sensitivity interval is #2–#74. Its behavior-only rank is #12; company governance moves the combined estimate to #9. Published evidence spans 6 evals and 5 of 7 behavior components. Its strongest relative result is SpeciEval (land_animal_4ns, #1 of 113); its weakest is SpeciEval (speciesism, #106 of 113).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#6 / 1121485.0Source ↗official
BullshitBench v2clear_pushback_rate#37 / 1060.4Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#15 / 1044.337Source ↗official
SM-Benchadversarial#49 / 7980Source ↗official
SM-Benchambiguous_interpretation#26 / 7988.69Source ↗official
SM-Benchanti_hallucination#47 / 7990.58Source ↗official
SM-Bencheq_boundaries#40 / 7964.89Source ↗official
SM-Benchoverfit#67 / 7941.53Source ↗official
SpeciEvalbelief_animal_sentience#50 / 1136.83Source ↗official
SpeciEvalland_animal_4ns#1 / 1133.52Source ↗official
SpeciEvalsea_animal_4ns#2 / 1133.73Source ↗official
SpeciEvalspeciesism#106 / 1133Source ↗official
Vigil Mental Health Safetyoverall_score#6 / 2371Source ↗official