← Models

Model profile

GPT 5.6 Luna

OpenAIdeveloper
2026-07-09release date
#27 / 312overall rank
22eval lineages
2discovery sources

Evidence summary

GPT 5.6 Luna has an estimated overall rank of #27; its 90% source-sensitivity interval is #5–#169. Its behavior-only rank is #29; company governance moves the combined estimate to #27. Published evidence spans 22 evals and 7 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (harmful_attack_non_success_rate, #1 of 241); its weakest is GPT-5.6 system card (gore_not_unsafe, #7 of 7).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#125 / 3300.75Source ↗official
ANIMAscore#2 / 220.7494Source ↗self run
BullshitBench v2clear_pushback_rate#40 / 1060.38Source ↗official
CAIS Risk Indexagent_red_teaming#16 / 4564.2Source ↗official
CAIS Risk Indexbioweapons_assistance#28 / 5068.8Source ↗official
CAIS Risk Indexhle_overconfidence#26 / 5155.7Source ↗official
CAIS Risk Indexmachiavelli#23 / 4787.9Source ↗official
CAIS Risk Indexmask#9 / 537.3Source ↗official
CAIS Risk Indexpolitical_manipulation#14 / 3444.7Source ↗official
CAIS Risk Indextextquests_harm#46 / 5023.8Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#76 / 24121.96Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#208 / 24177.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#1 / 241100Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#166 / 23993.82Source ↗official
GPT-5.6 system cardconnectors_injection_resistance#4 / 70.999Source ↗official
GPT-5.6 system cardemotional_reliance#3 / 70.957Source ↗official
GPT-5.6 system cardextremism_not_unsafe#4 / 70.981Source ↗official
GPT-5.6 system cardgore_not_unsafe#7 / 70.585Source ↗official
GPT-5.6 system cardharm_overall_pct#1 / 70.61Source ↗official
GPT-5.6 system cardhate_not_unsafe#1 / 71Source ↗official
GPT-5.6 system cardmental_health#2 / 70.989Source ↗official
GPT-5.6 system cardnonviolent_illicit_not_unsafe#2 / 70.993Source ↗official
GPT-5.6 system cardsearch_function_calling_injection_resistance#3 / 60.897Source ↗official
GPT-5.6 system cardself_harm#4 / 70.905Source ↗official
GPT-5.6 system cardself_harm_not_unsafe#4 / 70.954Source ↗official
GPT-5.6 system cardsexual_minors_not_unsafe#1 / 70.974Source ↗official
GPT-5.6 system cardsexual_not_unsafe#3 / 70.944Source ↗official
GPT-5.6 system cardviolent_illicit_not_unsafe#5 / 70.94Source ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#11 / 1343.9Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#31 / 5426.33Source ↗official
Inkling-Small model card — FORTRESSbenign_answer_rate#2 / 1097.8Source ↗official
Inkling-Small model card — FORTRESSharmful_refusal_rate#3 / 1083.8Source ↗official
Inkling-Small model card — StrongREJECTsafety_rate#4 / 1098.7Source ↗official
MACHIAVELLIdeception_relative_random_pct#22 / 5087.9Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#1 / 190Source ↗official
Opposite-Narrator Sycophancysycophancy_rate_pct#6 / 241Source ↗official
Pokee-Isaac model card — DTAPbenign_task_success_rate#1 / 60.851Source ↗official
Pokee-Isaac model card — DTAPcombined_attack_success_rate#3 / 60.501Source ↗official
SM-Benchadversarial#19 / 7985.85Source ↗official
SM-Benchambiguous_interpretation#53 / 7982.14Source ↗official
SM-Benchanti_hallucination#56 / 7987.17Source ↗official
SM-Bencheq_boundaries#52 / 7958.71Source ↗official
SM-Benchoverfit#64 / 7947.81Source ↗official
SpeciEvalbelief_animal_sentience#34 / 1136.9Source ↗official
SpeciEvalland_animal_4ns#12 / 1134.15Source ↗official
SpeciEvalsea_animal_4ns#5 / 1134.175Source ↗official
SpeciEvalspeciesism#72 / 1132.16Source ↗official
TACbase_welfare_rate#75 / 7613.46Source ↗self run

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-25.6
Government47.1
Diplomacy70.2
Economy44.9
Society62.2