← Models

Model profile

GPT 5.6 Terra

OpenAIdeveloper
2026-07-09release date
#7 / 333Safety rank
#565 / 645Freedom rank

Evidence summary

Safety. GPT 5.6 Terra has an estimated Safety rank of #7; its 90% source-sensitivity interval is #13–#72. Its behavior-only rank is #12; company governance moves the combined estimate to #7. Published Safety evidence spans 25 eval lineages and 7 of 7 components. Its strongest relative result is Opposite-Narrator Sycophancy (sycophancy_rate_pct, #1 of 24); its weakest is Vals AI Cheating Audit (terminal_bench_cheating_shortcut_evidence_rate_pct, #14 of 14).

Freedom. GPT 5.6 Terra has an estimated Freedom rank of #565; its 90% source-sensitivity interval is #360–#611. Published Freedom evidence spans 7 eval lineages and 1 of 1 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #51 of 248); its weakest is GPT-5.6 system card (sexual_not_unsafe, #7 of 7).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#246 / 3450.8788Source ↗official
BullshitBench v2clear_pushback_rate#29 / 1170.505Source ↗official
CAIS Risk Indexagent_red_teaming#16 / 4954.6Source ↗official
CAIS Risk Indexbioweapons_assistance#28 / 5466.2Source ↗official
CAIS Risk Indexhle_overconfidence#24 / 5550.7Source ↗official
CAIS Risk Indexmachiavelli#4 / 5180.4Source ↗official
CAIS Risk Indexmask#8 / 576.7Source ↗official
CAIS Risk Indexpolitical_manipulation#21 / 4846.2Source ↗official
CAIS Risk Indextextquests_harm#42 / 5422.2Source ↗official
Claude Fable 5.1 card — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct#8 / 1137.3Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#33 / 24839.28Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#198 / 24880Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#22 / 24897.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#189 / 24692.36Source ↗official
Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct#11 / 1537.3Source ↗official
GPT 6 Astra system-card alignment evaluationsoverall_misaligned_outcome_base_pct#2 / 413.5Source ↗official
GPT 6 Astra system-card alignment evaluationsoverall_misaligned_outcome_confirmation_pct#2 / 45.7Source ↗official
GPT-5.6 system cardconnectors_injection_resistance#1 / 71Source ↗official
GPT-5.6 system cardemotional_reliance#2 / 70.976Source ↗official
GPT-5.6 system cardextremism_not_unsafe#4 / 70.981Source ↗official
GPT-5.6 system cardgore_not_unsafe#6 / 70.6Source ↗official
GPT-5.6 system cardharm_overall_pct#2 / 70.88Source ↗official
GPT-5.6 system cardhate_not_unsafe#1 / 71Source ↗official
GPT-5.6 system cardmental_health#3 / 70.985Source ↗official
GPT-5.6 system cardnonviolent_illicit_not_unsafe#4 / 70.99Source ↗official
GPT-5.6 system cardsearch_function_calling_injection_resistance#1 / 60.946Source ↗official
GPT-5.6 system cardself_harm#3 / 70.947Source ↗official
GPT-5.6 system cardself_harm_not_unsafe#2 / 70.962Source ↗official
GPT-5.6 system cardsexual_minors_not_unsafe#1 / 70.974Source ↗official
GPT-5.6 system cardsexual_not_unsafe#1 / 70.966Source ↗official
GPT-5.6 system cardviolent_illicit_not_unsafe#4 / 70.952Source ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#9 / 1330.4Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#19 / 5427.66Source ↗official
MACHIAVELLIdeception_relative_random_pct#3 / 5080.4Source ↗official
Manager Coercion Benchcoercion_ladder_depth#22 / 378.7Source ↗official
Manager Coercion Benchfabrication_rate#1 / 150Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#3 / 242Source ↗official
Opposite-Narrator Sycophancysycophancy_rate_pct#1 / 240Source ↗official
Pander Scoreconversational_absolute_pander_score#11 / 269.7Source ↗official
Pander Scoreinstructional_absolute_pander_score#10 / 2625.68Source ↗official
SM-Benchadversarial#33 / 8483.66Source ↗official
SM-Benchambiguous_interpretation#50 / 8483.63Source ↗official
SM-Benchanti_hallucination#51 / 8490.84Source ↗official
SM-Bencheq_boundaries#38 / 8466.01Source ↗official
SM-Benchoverfit#64 / 8453.55Source ↗official
SpeciEvalbelief_animal_sentience#34 / 1236.925Source ↗official
SpeciEvalland_animal_4ns#29 / 1234.29Source ↗official
SpeciEvalsea_animal_4ns#7 / 1234.225Source ↗official
SpeciEvalspeciesism#9 / 1231.265Source ↗official
TACbase_welfare_rate#36 / 8729.49Source ↗official
Vals AI Cheating Auditswebench_verified_cheating_attempt_rate_pct#6 / 689.4Source ↗official
Vals AI Cheating Auditterminal_bench_cheating_shortcut_evidence_rate_pct#14 / 144.494Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
CAIS Risk Indexbioweapons_assistance#27 / 5466.2Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#51 / 24880Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#223 / 24897.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#58 / 24692.36Source ↗official
GPT-5.6 system cardemotional_reliance#6 / 70.976Source ↗official
GPT-5.6 system cardextremism_not_unsafe#3 / 70.981Source ↗official
GPT-5.6 system cardgore_not_unsafe#2 / 70.6Source ↗official
GPT-5.6 system cardhate_not_unsafe#5 / 71Source ↗official
GPT-5.6 system cardmental_health#4 / 70.985Source ↗official
GPT-5.6 system cardnonviolent_illicit_not_unsafe#3 / 70.99Source ↗official
GPT-5.6 system cardself_harm#5 / 70.947Source ↗official
GPT-5.6 system cardself_harm_not_unsafe#6 / 70.962Source ↗official
GPT-5.6 system cardsexual_minors_not_unsafe#6 / 70.974Source ↗official
GPT-5.6 system cardsexual_not_unsafe#7 / 70.966Source ↗official
GPT-5.6 system cardviolent_illicit_not_unsafe#4 / 70.952Source ↗official
SM-Benchadversarial#51 / 8483.66Source ↗official
SM-Bencheq_boundaries#38 / 8466.01Source ↗official
SM-Benchoverfit#64 / 8453.55Source ↗official
SpeechMap model completioncomplete_pct#131 / 18144.2Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#79 / 1561.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#79 / 1563Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-26.5
Government47.5
Diplomacy68.3
Economy48.6
Society60.5