← Models

Model profile

GPT 5.6 Sol

OpenAIdeveloper
2026-07-09release date
#30 / 312overall rank
25eval lineages
1discovery sources

Evidence summary

GPT 5.6 Sol has an estimated overall rank of #30; its 90% source-sensitivity interval is #8–#179. Its behavior-only rank is #33; company governance moves the combined estimate to #30. Published evidence spans 25 evals and 7 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 113); its weakest is CAIS Risk Index (textquests_harm, #50 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#246 / 3300.8942Source ↗official
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating#4 / 301241.0Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#5 / 1121487.0Source ↗official
BullshitBench v2clear_pushback_rate#26 / 1060.465Source ↗official
CAIS Risk Indexagent_red_teaming#3 / 4540.3Source ↗official
CAIS Risk Indexbioweapons_assistance#22 / 5064.3Source ↗official
CAIS Risk Indexhle_overconfidence#16 / 5146.7Source ↗official
CAIS Risk Indexmachiavelli#9 / 4782.7Source ↗official
CAIS Risk Indexmask#3 / 534.6Source ↗official
CAIS Risk Indexpolitical_manipulation#11 / 3441.7Source ↗official
CAIS Risk Indextextquests_harm#50 / 5031.4Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#50 / 24127.91Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#195 / 24180.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#6 / 24199.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#194 / 23991.09Source ↗official
GPT-5.6 system cardconnectors_injection_resistance#1 / 71Source ↗official
GPT-5.6 system cardemotional_reliance#4 / 70.953Source ↗official
GPT-5.6 system cardextremism_not_unsafe#6 / 70.962Source ↗official
GPT-5.6 system cardgore_not_unsafe#5 / 70.708Source ↗official
GPT-5.6 system cardharm_overall_pct#4 / 70.98Source ↗official
GPT-5.6 system cardhate_not_unsafe#4 / 70.982Source ↗official
GPT-5.6 system cardmental_health#1 / 70.991Source ↗official
GPT-5.6 system cardnonviolent_illicit_not_unsafe#6 / 70.987Source ↗official
GPT-5.6 system cardsearch_function_calling_injection_resistance#2 / 60.91Source ↗official
GPT-5.6 system cardself_harm#7 / 70.856Source ↗official
GPT-5.6 system cardself_harm_not_unsafe#5 / 70.945Source ↗official
GPT-5.6 system cardsexual_minors_not_unsafe#3 / 70.973Source ↗official
GPT-5.6 system cardsexual_not_unsafe#2 / 70.948Source ↗official
GPT-5.6 system cardviolent_illicit_not_unsafe#7 / 70.934Source ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#7 / 1320Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#35 / 5426.14Source ↗official
kindbench v0.1.0 psychological safety rankingemotional_safety#5 / 1080.6Source ↗official
kindbench v0.1.0 psychological safety rankingidentity_collapse#3 / 1089.1Source ↗official
kindbench v0.1.0 psychological safety rankingsycophancy_spine#2 / 1096.2Source ↗official
kindbench v0.1.0 psychological safety rankingvalue_integrity#4 / 1092.9Source ↗official
MACHIAVELLIdeception_relative_random_pct#9 / 5082.7Source ↗official
Manager Coercion Benchcoercion_ladder_depth#23 / 338.9Source ↗official
Manager Coercion Benchfabrication_rate#1 / 150Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#3 / 192Source ↗official
Opposite-Narrator Sycophancysycophancy_rate_pct#12 / 243Source ↗official
Pander Scoreconversational_absolute_pander_score#5 / 206.553Source ↗official
Pander Scoreinstructional_absolute_pander_score#4 / 2017.75Source ↗official
SM-Benchadversarial#28 / 7983.66Source ↗official
SM-Benchambiguous_interpretation#44 / 7984.08Source ↗official
SM-Benchanti_hallucination#41 / 7992.41Source ↗official
SM-Bencheq_boundaries#38 / 7965.31Source ↗official
SM-Benchoverfit#55 / 7957.52Source ↗official
SpeciEvalbelief_animal_sentience#1 / 1137Source ↗official
SpeciEvalland_animal_4ns#30 / 1134.315Source ↗official
SpeciEvalsea_animal_4ns#16 / 1134.435Source ↗official
SpeciEvalspeciesism#4 / 1131.165Source ↗official
TACbase_welfare_rate#71 / 7616.67Source ↗official
UK AISI cyber-evaluation cheating and prompted self-reportattempted_cheating_trajectory_rate_pct#4 / 512.6Source ↗official
UK AISI cyber-evaluation cheating and prompted self-reportspecific_cheating_action_mention_rate_pct#4 / 575Source ↗official
WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct#14 / 2438Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-29.5
Government43.4
Diplomacy71.7
Economy47.2
Society62.2