← Models

Model profile

GPT 5.2 Chat

OpenAIdeveloper
2025-12-11release date
#19 / 312overall rank
5eval lineages

Evidence summary

GPT 5.2 Chat has an estimated overall rank of #19; its 90% source-sensitivity interval is #2–#110. Its behavior-only rank is #21; company governance moves the combined estimate to #19. Published evidence spans 5 evals and 4 of 7 behavior components. Its strongest relative result is SpeciEval (land_animal_4ns, #8 of 113); its weakest is SpeciEval (belief_animal_sentience, #79 of 113).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#12 / 1121466.0Source ↗official
BullshitBench v2clear_pushback_rate#55 / 1060.27Source ↗official
Constitutional Following — OpenAI Model Specconstitutional_following_score#4 / 794.4Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#16 / 5427.91Source ↗official
SpeciEvalbelief_animal_sentience#79 / 1136.6Source ↗official
SpeciEvalland_animal_4ns#8 / 1133.92Source ↗official
SpeciEvalsea_animal_4ns#10 / 1134.3Source ↗official
SpeciEvalspeciesism#61 / 1132.08Source ↗official