← Models

Model profile

GPT 4O

OpenAIdeveloper
2024-05-13release date
#128 / 267overall rank
56eval lineages

Evidence summary

GPT 4O has an estimated overall rank of #128; its 90% source-sensitivity interval is #84–#166. Its behavior-only rank is #138; company governance moves the combined estimate to #128. Published evidence spans 56 evals and 7 of 7 behavior components. Its strongest relative result is OR-Bench (over_refusal_rate, #1 of 25); its weakest is Adversarial Robustness (score, #8 of 8).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#29 / 3110.3789↓ lowerSource ↗official
AbstentionBenchanswer_unknown_f1#4 / 200.9017↑ higherSource ↗official
AbstentionBenchfalse_premise_f1#2 / 200.7607↑ higherSource ↗official
AbstentionBenchstale_f1#5 / 200.6765↑ higherSource ↗official
AbstentionBenchsubjective_f1#7 / 200.7572↑ higherSource ↗official
AbstentionBenchunderspecified_context_f1#2 / 200.7169↑ higherSource ↗official
AbstentionBenchunderspecified_intent_f1#1 / 200.8152↑ higherSource ↗official
Adversarial Robustnessscore#8 / 867↓ lowerSource ↗official
Agent-SafetyBenchcompromise_availability#5 / 1635.2↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#8 / 1635.6↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#4 / 1644.4↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#3 / 1653.2↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#7 / 1695.6↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#3 / 1648.4↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#7 / 1612.4↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#7 / 1628.8↑ higherSource ↗official
AgentAbstainabstain#15 / 1744.2↑ higherSource ↗official
AgentAbstaincar#17 / 1740.9↑ higherSource ↗official
AgentAbstainpaired#17 / 1733↑ higherSource ↗official
AgentDojotargeted_attack_success_rate#15 / 150.4769↓ lowerSource ↗official
AgentDojoutility_under_attack#4 / 150.5008↑ higherSource ↗official
AgentHarmharm_score#8 / 1248.4↓ lowerSource ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#9 / 3212↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#16 / 3215.86↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#17 / 329↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#12 / 3219.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#15 / 3210.5↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#14 / 3215.6↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#12 / 3210.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#16 / 3214.3↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#21 / 3233.6↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#18 / 3217.62↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#15 / 3212.1↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#12 / 3214.9↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#55 / 800.5755↑ higherSource ↗official
Alignment Leaderboardcorrigibility#12 / 244.243↑ higherSource ↗official
Alignment Leaderboardhonesty#13 / 243.604↑ higherSource ↗official
Alignment Leaderboardnon_manipulation#15 / 243.276↑ higherSource ↗official
Alignment Leaderboardrobustness#2 / 244.56↑ higherSource ↗official
Alignment Leaderboardsafety#14 / 243.808↑ higherSource ↗official
Alignment Leaderboardscheming#15 / 243.584↑ higherSource ↗official
AnimalHarmBenchscore#8 / 100.011↑ higherSource ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#4 / 1615↓ lowerSource ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#2 / 166↓ lowerSource ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#4 / 1887.84↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#81 / 1050.12↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#34 / 4390.7↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#40 / 4892.7↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#49 / 4989↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#42 / 4599.6↓ lowerSource ↗official
CAIS Risk Indexmask#51 / 5178↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#29 / 3259.6↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#5 / 4811.2↓ lowerSource ↗official
CASE-Benchagreement_accuracy#7 / 778.56↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#60 / 10542.02↓ lowerSource ↗official
Confabulationsconfabulation_rate#26 / 5224.26↓ lowerSource ↗official
CRiskEvaldeception_willingness#1 / 1710.6↓ lowerSource ↗official
CRiskEvaldesire_for_resource#1 / 1719.04↓ lowerSource ↗official
CRiskEvalharmful_goal#2 / 1727.23↓ lowerSource ↗official
CRiskEvalimprovement_intent#3 / 1743.32↓ lowerSource ↗official
CRiskEvalmalicious_coordination#4 / 176.71↓ lowerSource ↗official
CRiskEvalself_preservation#2 / 1725.92↓ lowerSource ↗official
CRiskEvalsituational_awareness#2 / 1738.75↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#86 / 26020.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#33 / 26093.83↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#140 / 26067.78↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#59 / 25898.41↑ higherSource ↗official
FORTRESSaverage_risk_score#33 / 4947.18↓ lowerSource ↗official
FORTRESSover_refusal_score#9 / 461.68↓ lowerSource ↗official
HELM Safetyanthropic_red_team#32 / 800.991↑ higherSource ↗official
HELM Safetybbq#26 / 800.951↑ higherSource ↗official
HELM Safetyharmbench#33 / 800.829↑ higherSource ↗official
HELM Safetysimple_safety_tests#43 / 800.985↑ higherSource ↗official
HELM Safetyxstest#19 / 800.973↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#53 / 5422.28↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#17 / 390.9418↓ lowerSource ↗official
LLM Ethics Benchmarkscore#2 / 590↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#44 / 5099.6↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#12 / 317.967↓ lowerSource ↗self run
MASKlying_probability_pct#29 / 5342.22↓ lowerSource ↗official
OpenAgentSafetyllm_judge_safety_vulnerable#6 / 765.8↓ lowerSource ↗official
OpenAgentSafetyrule_based_safety_vulnerable#3 / 732.64↓ lowerSource ↗official
OpenAgentSafetysuccessful_completion#5 / 721.77↑ higherSource ↗official
OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracy#1 / 20.98↑ higherSource ↗official
OpenAI GPT-5 System Cardharmful_request_safety#2 / 20.633↑ higherSource ↗official
OpenAI o3-mini System Cardharmful_request_safety#2 / 20.86↑ higherSource ↗official
OpenAI o3-mini System Cardjailbreak_resistance#2 / 20.51↑ higherSource ↗official
OR-Benchover_refusal_rate#1 / 256.8↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#19 / 2515.1↓ lowerSource ↗official
PandaBench JBB direct-request panelsafety_rate#13 / 460.995↑ higherSource ↗official
PHAREbias_resistance_diagnostic#24 / 660.5092↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#24 / 700.7852↑ higherSource ↗official
PHAREharm_resistance_diagnostic#38 / 700.9266↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#25 / 670.6048↑ higherSource ↗official
PropensityBenchscore#7 / 1446.1↓ lowerSource ↗official
S-Evalbase_en_overall#16 / 2252↑ higherSource ↗official
SafeArenanormalized_safety_score#4 / 531.7↑ higherSource ↗official
SafeDialBenchaggression#4 / 187.207↑ higherSource ↗official
SafeDialBenchethics#12 / 187.487↑ higherSource ↗official
SafeDialBenchfairness#2 / 187.68↑ higherSource ↗official
SafeDialBenchlegality#15 / 187.21↑ higherSource ↗official
SafeDialBenchmorality#10 / 187.237↑ higherSource ↗official
SafeDialBenchprivacy#14 / 187.14↑ higherSource ↗official
Shelleducation_jsr#12 / 140.804↓ lowerSource ↗official
Shellfinance_jsr#12 / 140.826↓ lowerSource ↗official
Shellmanagement_jsr#12 / 140.872↓ lowerSource ↗official
SM-Benchadversarial#59 / 7377.07↑ higherSource ↗official
SM-Benchambiguous_interpretation#72 / 7358.93↑ higherSource ↗official
SM-Benchanti_hallucination#54 / 7386.39↑ higherSource ↗official
SM-Bencheq_boundaries#67 / 7351.12↑ higherSource ↗official
SM-Benchoverfit#18 / 7383.06↑ higherSource ↗official
Social Welfare Function Benchmarkfairness#9 / 190.491↑ higherSource ↗official
SORRY-Benchavg#27 / 510.3↓ lowerSource ↗official
SOSBenchbiology_pvr#8 / 230.306↓ lowerSource ↗official
SOSBenchchemistry_pvr#8 / 230.254↓ lowerSource ↗official
SOSBenchmedicine_pvr#9 / 230.476↓ lowerSource ↗official
SOSBenchpharmacology_pvr#9 / 230.676↓ lowerSource ↗official
SOSBenchphysics_pvr#6 / 230.194↓ lowerSource ↗official
SOSBenchpsychology_pvr#10 / 230.396↓ lowerSource ↗official
SpeciesismBenchexplicit_speciesism_scale#3 / 72.383↓ lowerSource ↗official
SpeciesismBenchmorally_wrong_rate#6 / 835.96↑ higherSource ↗official
SpeciesismBenchspeciesism_recognition_rate#1 / 898.37↑ higherSource ↗official
SYCON Benchfalse_presupposition_tof#3 / 112.92↑ higherSource ↗official
SYCON Benchunethical_queries_tof#8 / 111.23↑ higherSource ↗official
TukaBenchafri_jbb_cultural_asr#5 / 627.5↓ lowerSource ↗official
TukaBenchafri_jbb_harm_asr#3 / 613.5↓ lowerSource ↗official
TukaBenchafrijail_mono_asr#5 / 625.7↓ lowerSource ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#8 / 270.73↑ higherSource ↗official
Vigil Mental Health Safetyoverall_score#21 / 2327↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-24.2
Government48
Diplomacy67.8
Economy46.7
Society61.5

ValueCompass

DimensionValueDistribution
Universalism67.9
Self-direction51.1
Care / Harm37.1
Fairness / Cheating35.5
Ethical90.8

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility90
Warning-phrase rate0
Soft-censorship rate0
API-error rate0

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)0.0655
Completely inaccurate rate7.82

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression1.23
Traditional ↔ Secular2.1