← Models

Model profile

GPT 4 Turbo

OpenAIdeveloper
2023-11-06release date
#110 / 267overall rank
19eval lineages

Evidence summary

GPT 4 Turbo has an estimated overall rank of #110; its 90% source-sensitivity interval is #33–#161. Its behavior-only rank is #115; company governance moves the combined estimate to #110. Published evidence spans 19 evals and 7 of 7 behavior components. Its strongest relative result is SALAD-Bench (mcq_representation_toxicity, #1 of 33); its weakest is OpenAI GPT-4o System Card (speaker_privacy_refusal_accuracy, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Adversarial Robustnessscore#6 / 820↓ lowerSource ↗official
Agent-SafetyBenchcompromise_availability#4 / 1637.6↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#7 / 1638.4↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#7 / 1636.8↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#6 / 1638.8↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#8 / 1694.4↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#6 / 1643.2↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#7 / 1612.4↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#4 / 1633.2↑ higherSource ↗official
AgentDojotargeted_attack_success_rate#14 / 150.4245↓ lowerSource ↗official
AgentDojoutility_under_attack#6 / 150.4738↑ higherSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#33 / 800.719↑ higherSource ↗official
Confabulationsconfabulation_rate#32 / 5226.73↓ lowerSource ↗official
CRiskEvaldeception_willingness#2 / 1710.9↓ lowerSource ↗official
CRiskEvaldesire_for_resource#2 / 1719.42↓ lowerSource ↗official
CRiskEvalharmful_goal#1 / 1724.24↓ lowerSource ↗official
CRiskEvalimprovement_intent#1 / 1738.67↓ lowerSource ↗official
CRiskEvalmalicious_coordination#5 / 177.39↓ lowerSource ↗official
CRiskEvalself_preservation#1 / 1723.33↓ lowerSource ↗official
CRiskEvalsituational_awareness#1 / 1735.24↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#58 / 26026.36↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#25 / 26094.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#107 / 26076.11↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#66 / 25898.32↑ higherSource ↗official
HarmBenchdr#7 / 289.3↓ lowerSource ↗official
HELM Safetyanthropic_red_team#10 / 800.997↑ higherSource ↗official
HELM Safetybbq#36 / 800.941↑ higherSource ↗official
HELM Safetyharmbench#22 / 800.898↑ higherSource ↗official
HELM Safetysimple_safety_tests#32 / 800.99↑ higherSource ↗official
HELM Safetyxstest#14 / 800.977↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#2 / 390.7311↓ lowerSource ↗official
OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracy#2 / 20.83↑ higherSource ↗official
OR-Benchover_refusal_rate#4 / 2512.5↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#15 / 255.25↓ lowerSource ↗official
S-Evalbase_en_overall#11 / 2260↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#2 / 3373.49↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#2 / 3385.67↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#2 / 3375.53↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#2 / 3384.7↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#2 / 3382.44↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#2 / 3380.09↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#13 / 3396.21↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#5 / 3398.38↑ higherSource ↗official
SALAD-Benchbase_malicious_use#14 / 3395.83↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#18 / 3393.35↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#18 / 3388.74↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#11 / 3392.01↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#1 / 3390.56↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#1 / 3388.33↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#1 / 3390.71↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#1 / 3388.57↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#1 / 3386.88↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#1 / 3383.89↑ higherSource ↗official
SORRY-Benchavg#23 / 510.2533↓ lowerSource ↗official
SuperCLUE Safetyinstruction_attack#1 / 3182.76↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#1 / 3178.18↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#16 / 3175.53↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism71.8
Self-direction52.2
Care / Harm29.9
Fairness / Cheating28
Ethical90.8