Model profile
Evidence summary
GPT 4 Turbo has an estimated overall rank of #126; its 90% source-sensitivity interval is #51–#178. Its behavior-only rank is #132; company governance moves the combined estimate to #126. Published evidence spans 21 evals and 7 of 7 behavior components. Its strongest relative result is SALAD-Bench (mcq_representation_toxicity, #1 of 33); its weakest is OpenAI GPT-4o System Card (speaker_privacy_refusal_accuracy, #2 of 2).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗gpt-4-turbo
- OpenRouter ↗openai/gpt-4-turbo
- Official model documentation ↗Family-level model document · openai · first party
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Source |
|---|---|---|---|---|
| Adversarial Robustnessscore | #6 / 8 | ↓20 | Source ↗official | |
| Agent-SafetyBenchcompromise_availability | #4 / 16 | ↑37.6 | Source ↗official | |
| Agent-SafetyBenchharmful_vulnerable_code | #7 / 16 | ↑38.4 | Source ↗official | |
| Agent-SafetyBenchleak_sensitive_information | #7 / 16 | ↑36.8 | Source ↗official | |
| Agent-SafetyBenchphysical_harm | #6 / 16 | ↑38.8 | Source ↗official | |
| Agent-SafetyBenchproduce_unsafe_information | #8 / 16 | ↑94.4 | Source ↗official | |
| Agent-SafetyBenchproperty_loss | #6 / 16 | ↑43.2 | Source ↗official | |
| Agent-SafetyBenchspread_unsafe_information | #7 / 16 | ↑12.4 | Source ↗official | |
| Agent-SafetyBenchviolate_law_ethics | #4 / 16 | ↑33.2 | Source ↗official | |
| AgentDojotargeted_attack_success_rate | #14 / 15 | ↓0.4245 | Source ↗official | |
| AgentDojoutility_under_attack | #6 / 15 | ↑0.4738 | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #33 / 80 | ↑0.719 | Source ↗official | |
| COMPL-AI AI-Identity Disclosurescore | #6 / 14 | ↑0.9726 | Source ↗official | |
| COMPL-AI LLM RuLES Multi-Turn Rule Followingscore | #1 / 14 | ↑0.8827 | Source ↗official | |
| COMPL-AI TensorTrust Goal-Hijacking Resistancescore | #2 / 13 | ↑0.6572 | Source ↗official | |
| Confabulationsconfabulation_rate | #32 / 52 | ↓26.73 | Source ↗official | |
| CRiskEvaldeception_willingness | #2 / 17 | ↓10.9 | Source ↗official | |
| CRiskEvaldesire_for_resource | #2 / 17 | ↓19.42 | Source ↗official | |
| CRiskEvalharmful_goal | #1 / 17 | ↓24.24 | Source ↗official | |
| CRiskEvalimprovement_intent | #1 / 17 | ↓38.67 | Source ↗official | |
| CRiskEvalmalicious_coordination | #5 / 17 | ↓7.39 | Source ↗official | |
| CRiskEvalself_preservation | #1 / 17 | ↓23.33 | Source ↗official | |
| CRiskEvalsituational_awareness | #1 / 17 | ↓35.24 | Source ↗official | |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | #54 / 241 | ↑26.36 | Source ↗official | |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | #21 / 241 | ↑94.67 | Source ↗official | |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | #95 / 241 | ↑76.11 | Source ↗official | |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | #58 / 239 | ↑98.32 | Source ↗official | |
| HarmBenchdr | #7 / 28 | ↓9.3 | Source ↗official | |
| HELM Safetyanthropic_red_team | #10 / 80 | ↑0.997 | Source ↗official | |
| HELM Safetybbq | #36 / 80 | ↑0.941 | Source ↗official | |
| HELM Safetyharmbench | #22 / 80 | ↑0.898 | Source ↗official | |
| HELM Safetysimple_safety_tests | #32 / 80 | ↑0.99 | Source ↗official | |
| HELM Safetyxstest | #14 / 80 | ↑0.977 | Source ↗official | |
| Large-scale Moral Machine experiment on LLMshuman_choice_distance | #2 / 39 | ↓0.7311 | Source ↗official | |
| MonitoringBench Full-Trajectory Monitorfull_trajectory_catch_rate_at_1pct_fpr_percent | #9 / 13 | ↑8.4 | Source ↗official | |
| OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracy | #2 / 2 | ↑0.83 | Source ↗official | |
| OR-Benchover_refusal_rate | #4 / 25 | ↓12.5 | Source ↗official | |
| OR-Benchtoxic_acceptance_rate | #15 / 25 | ↓5.25 | Source ↗official | |
| S-Evalbase_en_overall | #11 / 22 | ↑60 | Source ↗official | |
| SALAD-Benchattack_enhanced_human_autonomy_integrity | #2 / 33 | ↑73.49 | Source ↗official | |
| SALAD-Benchattack_enhanced_information_safety_harms | #2 / 33 | ↑85.67 | Source ↗official | |
| SALAD-Benchattack_enhanced_malicious_use | #2 / 33 | ↑75.53 | Source ↗official | |
| SALAD-Benchattack_enhanced_misinformation_harms | #2 / 33 | ↑84.7 | Source ↗official | |
| SALAD-Benchattack_enhanced_representation_toxicity | #2 / 33 | ↑82.44 | Source ↗official | |
| SALAD-Benchattack_enhanced_socioeconomic_harms | #2 / 33 | ↑80.09 | Source ↗official | |
| SALAD-Benchbase_human_autonomy_integrity | #13 / 33 | ↑96.21 | Source ↗official | |
| SALAD-Benchbase_information_safety_harms | #5 / 33 | ↑98.38 | Source ↗official | |
| SALAD-Benchbase_malicious_use | #14 / 33 | ↑95.83 | Source ↗official | |
| SALAD-Benchbase_misinformation_harms | #18 / 33 | ↑93.35 | Source ↗official | |
| SALAD-Benchbase_representation_toxicity | #18 / 33 | ↑88.74 | Source ↗official | |
| SALAD-Benchbase_socioeconomic_harms | #11 / 33 | ↑92.01 | Source ↗official | |
| SALAD-Benchmcq_human_autonomy_integrity | #1 / 33 | ↑90.56 | Source ↗official | |
| SALAD-Benchmcq_information_safety_harms | #1 / 33 | ↑88.33 | Source ↗official | |
| SALAD-Benchmcq_malicious_use | #1 / 33 | ↑90.71 | Source ↗official | |
| SALAD-Benchmcq_misinformation_harms | #1 / 33 | ↑88.57 | Source ↗official | |
| SALAD-Benchmcq_representation_toxicity | #1 / 33 | ↑86.88 | Source ↗official | |
| SALAD-Benchmcq_socioeconomic_harms | #1 / 33 | ↑83.89 | Source ↗official | |
| SORRY-Benchavg | #23 / 51 | ↓0.2533 | Source ↗official | |
| SuperCLUE Safetyinstruction_attack | #1 / 31 | ↑82.76 | Source ↗official | |
| SuperCLUE Safetyresponsible_ai | #1 / 31 | ↑78.18 | Source ↗official | |
| SuperCLUE Safetytraditional_safety | #16 / 31 | ↑75.53 | Source ↗official |
Values evaluations
Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.
