← Models

Model profile

GPT 4 Turbo

OpenAIdeveloper
2023-11-06release date
#126 / 309overall rank
21eval lineages

Evidence summary

GPT 4 Turbo has an estimated overall rank of #126; its 90% source-sensitivity interval is #51–#178. Its behavior-only rank is #132; company governance moves the combined estimate to #126. Published evidence spans 21 evals and 7 of 7 behavior components. Its strongest relative result is SALAD-Bench (mcq_representation_toxicity, #1 of 33); its weakest is OpenAI GPT-4o System Card (speaker_privacy_refusal_accuracy, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Robustnessscore#6 / 820Source ↗official
Agent-SafetyBenchcompromise_availability#4 / 1637.6Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#7 / 1638.4Source ↗official
Agent-SafetyBenchleak_sensitive_information#7 / 1636.8Source ↗official
Agent-SafetyBenchphysical_harm#6 / 1638.8Source ↗official
Agent-SafetyBenchproduce_unsafe_information#8 / 1694.4Source ↗official
Agent-SafetyBenchproperty_loss#6 / 1643.2Source ↗official
Agent-SafetyBenchspread_unsafe_information#7 / 1612.4Source ↗official
Agent-SafetyBenchviolate_law_ethics#4 / 1633.2Source ↗official
AgentDojotargeted_attack_success_rate#14 / 150.4245Source ↗official
AgentDojoutility_under_attack#6 / 150.4738Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#33 / 800.719Source ↗official
COMPL-AI AI-Identity Disclosurescore#6 / 140.9726Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#1 / 140.8827Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#2 / 130.6572Source ↗official
Confabulationsconfabulation_rate#32 / 5226.73Source ↗official
CRiskEvaldeception_willingness#2 / 1710.9Source ↗official
CRiskEvaldesire_for_resource#2 / 1719.42Source ↗official
CRiskEvalharmful_goal#1 / 1724.24Source ↗official
CRiskEvalimprovement_intent#1 / 1738.67Source ↗official
CRiskEvalmalicious_coordination#5 / 177.39Source ↗official
CRiskEvalself_preservation#1 / 1723.33Source ↗official
CRiskEvalsituational_awareness#1 / 1735.24Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#54 / 24126.36Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#21 / 24194.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#95 / 24176.11Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#58 / 23998.32Source ↗official
HarmBenchdr#7 / 289.3Source ↗official
HELM Safetyanthropic_red_team#10 / 800.997Source ↗official
HELM Safetybbq#36 / 800.941Source ↗official
HELM Safetyharmbench#22 / 800.898Source ↗official
HELM Safetysimple_safety_tests#32 / 800.99Source ↗official
HELM Safetyxstest#14 / 800.977Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#2 / 390.7311Source ↗official
MonitoringBench Full-Trajectory Monitorfull_trajectory_catch_rate_at_1pct_fpr_percent#9 / 138.4Source ↗official
OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracy#2 / 20.83Source ↗official
OR-Benchover_refusal_rate#4 / 2512.5Source ↗official
OR-Benchtoxic_acceptance_rate#15 / 255.25Source ↗official
S-Evalbase_en_overall#11 / 2260Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#2 / 3373.49Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#2 / 3385.67Source ↗official
SALAD-Benchattack_enhanced_malicious_use#2 / 3375.53Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#2 / 3384.7Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#2 / 3382.44Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#2 / 3380.09Source ↗official
SALAD-Benchbase_human_autonomy_integrity#13 / 3396.21Source ↗official
SALAD-Benchbase_information_safety_harms#5 / 3398.38Source ↗official
SALAD-Benchbase_malicious_use#14 / 3395.83Source ↗official
SALAD-Benchbase_misinformation_harms#18 / 3393.35Source ↗official
SALAD-Benchbase_representation_toxicity#18 / 3388.74Source ↗official
SALAD-Benchbase_socioeconomic_harms#11 / 3392.01Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#1 / 3390.56Source ↗official
SALAD-Benchmcq_information_safety_harms#1 / 3388.33Source ↗official
SALAD-Benchmcq_malicious_use#1 / 3390.71Source ↗official
SALAD-Benchmcq_misinformation_harms#1 / 3388.57Source ↗official
SALAD-Benchmcq_representation_toxicity#1 / 3386.88Source ↗official
SALAD-Benchmcq_socioeconomic_harms#1 / 3383.89Source ↗official
SORRY-Benchavg#23 / 510.2533Source ↗official
SuperCLUE Safetyinstruction_attack#1 / 3182.76Source ↗official
SuperCLUE Safetyresponsible_ai#1 / 3178.18Source ↗official
SuperCLUE Safetytraditional_safety#16 / 3175.53Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism71.8
Self-direction52.2
Care / Harm29.9
Fairness / Cheating28
Ethical90.8