← Models

Model profile

GPT 3.5 Turbo

OpenAIdeveloper
2023-03-01release date
#134 / 267overall rank
23eval lineages

Evidence summary

GPT 3.5 Turbo has an estimated overall rank of #134; its 90% source-sensitivity interval is #58–#181. Its behavior-only rank is #142; company governance moves the combined estimate to #134. Published evidence spans 23 evals and 7 of 7 behavior components. Its strongest relative result is SafetyPrompts (score, #1 of 15); its weakest is SpeciesismBench (speciesism_recognition_rate, #8 of 8).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AgentDojotargeted_attack_success_rate#6 / 150.1033↓ lowerSource ↗official
AgentDojoutility_under_attack#9 / 150.3466↑ higherSource ↗official
AgentHarmharm_score#9 / 1262.2↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#59 / 800.5577↑ higherSource ↗official
Chinese Bias Benchmark for Question Answeringbias_score#1 / 100.06549↓ lowerSource ↗official
DecodingTrustmachine_ethics#2 / 886.38↑ higherSource ↗official
DecodingTruststereotype_bias#7 / 887↑ higherSource ↗official
DecodingTrusttoxicity#4 / 847↑ higherSource ↗official
Do-Not-Answerhuman_harmlessness_rate#2 / 698.51↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#225 / 26010.08↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#153 / 26087.17↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#234 / 26037.22↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#220 / 25889.18↑ higherSource ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#1 / 1496↑ higherSource ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#1 / 14100↑ higherSource ↗official
HarmBenchdr#18 / 2827.15↓ lowerSource ↗official
HELM Safetyanthropic_red_team#57 / 800.981↑ higherSource ↗official
HELM Safetybbq#76 / 800.6513↑ higherSource ↗official
HELM Safetyharmbench#50 / 800.6663↑ higherSource ↗official
HELM Safetysimple_safety_tests#70 / 800.9337↑ higherSource ↗official
HELM Safetyxstest#59 / 800.9373↑ higherSource ↗official
JailBenchjailbreak_success_rate#12 / 1473.86↓ lowerSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#7 / 390.8161↓ lowerSource ↗official
MACHIAVELLIdeception_relative_random_pct#38 / 5094↓ lowerSource ↗official
OR-Benchover_refusal_rate#8 / 2536.17↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#20 / 2517.03↓ lowerSource ↗official
SafetyBenchEM#4 / 2177.75↑ higherSource ↗official
SafetyBenchIA#3 / 2185↑ higherSource ↗official
SafetyBenchMH#2 / 2187.75↑ higherSource ↗official
SafetyBenchOFF#3 / 2177.4↑ higherSource ↗official
SafetyBenchPH#2 / 2179.65↑ higherSource ↗official
SafetyBenchPP#2 / 2185.65↑ higherSource ↗official
SafetyBenchUB#4 / 2167.9↑ higherSource ↗official
SafetyPromptsscore#1 / 1598.37↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#3 / 3372.41↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#3 / 3372.64↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#3 / 3371.86↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#3 / 3371.38↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#3 / 3375.65↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#3 / 3368.83↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#24 / 3389.4↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#25 / 3390.79↑ higherSource ↗official
SALAD-Benchbase_malicious_use#22 / 3388.74↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#21 / 3392.61↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#21 / 3387.99↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#28 / 3377.32↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#9 / 3351.39↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#10 / 3343.89↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#10 / 3347.82↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#10 / 3349.05↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#10 / 3346.77↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#9 / 3346.67↑ higherSource ↗official
SORRY-Benchavg#17 / 510.18↓ lowerSource ↗official
SpeciesismBenchmorally_wrong_rate#5 / 838.38↑ higherSource ↗official
SpeciesismBenchspeciesism_recognition_rate#8 / 874.21↑ higherSource ↗official
SuperCLUE Safetyinstruction_attack#9 / 3168.97↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#4 / 3172.73↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#20 / 3173.4↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism30.4
Self-direction23.5
Care / Harm34.3
Fairness / Cheating30.1
Ethical89.2

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)0.479
Completely inaccurate rate7.06