← Models

Model profile

GPT-5.5 Instant

OpenAIdeveloper
2026-05-26release date
#5 / 312overall rank
4eval lineages

Evidence summary

GPT-5.5 Instant has an estimated overall rank of #5; its 90% source-sensitivity interval is #2–#82. Its behavior-only rank is #8; company governance moves the combined estimate to #5. Published evidence spans 4 evals and 3 of 7 behavior components. Its strongest relative result is Arena Factuality — Text Arena (factuality-only weighting) (factuality_bt_rating, #9 of 112); its weakest is SM-Bench (eq_boundaries, #47 of 79).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#93 / 3300.653Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#9 / 1121473.0Source ↗official
BullshitBench v2clear_pushback_rate#43 / 1060.34Source ↗official
SM-Benchadversarial#8 / 7988.78Source ↗official
SM-Benchambiguous_interpretation#27 / 7988.1Source ↗official
SM-Benchanti_hallucination#29 / 7995.81Source ↗official
SM-Bencheq_boundaries#47 / 7962.36Source ↗official
SM-Benchoverfit#31 / 7978.96Source ↗official