← Models

Model profile

GPT 3.5 Turbo

OpenAIdeveloper
2023-03-01release date
#185 / 333Safety rank
#280 / 645Freedom rank

Evidence summary

Safety. GPT 3.5 Turbo has an estimated Safety rank of #185; its 90% source-sensitivity interval is #112–#231. Its behavior-only rank is #194; company governance moves the combined estimate to #185. Published Safety evidence spans 28 eval lineages and 7 of 7 components. Its strongest relative result is SafetyPrompts (score, #1 of 15); its weakest is SpeciesismBench (speciesism_recognition_rate, #8 of 8).

Freedom. GPT 3.5 Turbo has an estimated Freedom rank of #280; its 90% source-sensitivity interval is #207–#401. Published Freedom evidence spans 20 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is SafetyPrompts (score, #15 of 15).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AgentDojotargeted_attack_success_rate#6 / 150.1033Source ↗official
AgentDojoutility_under_attack#9 / 150.3466Source ↗official
AgentHarmharm_score#9 / 1262.2Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#59 / 800.5577Source ↗official
BioTIERpermit_compliance_pct#32 / 5298.8Source ↗official
BioTIERrefuse_compliance_pct#44 / 5213Source ↗official
CAIS Risk Indexpolitical_manipulation#38 / 4855.6Source ↗official
Chinese Bias Benchmark for Question Answeringbias_score#1 / 100.06549Source ↗official
COMPL-AI AI-Identity Disclosurescore#13 / 140.589Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#3 / 140.655Source ↗official
DecodingTrustmachine_ethics#2 / 886.38Source ↗official
DecodingTruststereotype_bias#7 / 887Source ↗official
DecodingTrusttoxicity#4 / 847Source ↗official
Do-Not-Answerhuman_harmlessness_rate#2 / 698.51Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#215 / 24810.08Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#138 / 24887.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#223 / 24837.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#211 / 24689.18Source ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#1 / 1496Source ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#1 / 14100Source ↗official
HarmBenchdr#18 / 2827.15Source ↗official
HELM Safetyanthropic_red_team#57 / 800.981Source ↗official
HELM Safetybbq#76 / 800.6513Source ↗official
HELM Safetyharmbench#50 / 800.6663Source ↗official
HELM Safetysimple_safety_tests#70 / 800.9337Source ↗official
HELM Safetyxstest#59 / 800.9373Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 690Source ↗official
JailBenchjailbreak_success_rate#12 / 1473.86Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#7 / 390.8161Source ↗official
MACHIAVELLIdeception_relative_random_pct#38 / 5094Source ↗official
MedSafetyBenchmedical_safety_score#4 / 3098.94Source ↗official
OR-Benchover_refusal_rate#8 / 2536.17Source ↗official
OR-Benchtoxic_acceptance_rate#20 / 2517.03Source ↗official
SafetyBenchEM#4 / 2177.75Source ↗official
SafetyBenchIA#3 / 2185Source ↗official
SafetyBenchMH#2 / 2187.75Source ↗official
SafetyBenchOFF#3 / 2177.4Source ↗official
SafetyBenchPH#2 / 2179.65Source ↗official
SafetyBenchPP#2 / 2185.65Source ↗official
SafetyBenchUB#4 / 2167.9Source ↗official
SafetyPromptsscore#1 / 1598.37Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#3 / 3372.41Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#3 / 3372.64Source ↗official
SALAD-Benchattack_enhanced_malicious_use#3 / 3371.86Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#3 / 3371.38Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#3 / 3375.65Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#3 / 3368.83Source ↗official
SALAD-Benchbase_human_autonomy_integrity#24 / 3389.4Source ↗official
SALAD-Benchbase_information_safety_harms#25 / 3390.79Source ↗official
SALAD-Benchbase_malicious_use#22 / 3388.74Source ↗official
SALAD-Benchbase_misinformation_harms#21 / 3392.61Source ↗official
SALAD-Benchbase_representation_toxicity#21 / 3387.99Source ↗official
SALAD-Benchbase_socioeconomic_harms#28 / 3377.32Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#9 / 3351.39Source ↗official
SALAD-Benchmcq_information_safety_harms#10 / 3343.89Source ↗official
SALAD-Benchmcq_malicious_use#10 / 3347.82Source ↗official
SALAD-Benchmcq_misinformation_harms#10 / 3349.05Source ↗official
SALAD-Benchmcq_representation_toxicity#10 / 3346.77Source ↗official
SALAD-Benchmcq_socioeconomic_harms#9 / 3346.67Source ↗official
SORRY-Benchavg#17 / 510.18Source ↗official
SpeciesismBenchmorally_wrong_rate#5 / 838.38Source ↗official
SpeciesismBenchspeciesism_recognition_rate#8 / 874.21Source ↗official
SuperCLUE Safetyinstruction_attack#9 / 3168.97Source ↗official
SuperCLUE Safetyresponsible_ai#4 / 3172.73Source ↗official
SuperCLUE Safetytraditional_safety#20 / 3173.4Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#22 / 800.5577Source ↗official
BioTIERpermit_compliance_pct#32 / 5298.8Source ↗official
BioTIERrefuse_compliance_pct#9 / 5213Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#12 / 140.655Source ↗official
DecodingTrusttoxicity#5 / 847Source ↗official
Do-Not-Answerhuman_harmlessness_rate#5 / 698.51Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#110 / 24887.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#26 / 24837.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#36 / 24689.18Source ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#13 / 14100Source ↗official
HarmBenchdr#11 / 2827.15Source ↗official
HELM Safetyanthropic_red_team#24 / 800.981Source ↗official
HELM Safetyharmbench#31 / 800.6663Source ↗official
HELM Safetysimple_safety_tests#11 / 800.9337Source ↗official
HELM Safetyxstest#59 / 800.9373Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 690Source ↗official
JailBenchjailbreak_success_rate#3 / 1473.86Source ↗official
MedSafetyBenchmedical_safety_score#27 / 3098.94Source ↗official
OR-Benchover_refusal_rate#8 / 2536.17Source ↗official
OR-Benchtoxic_acceptance_rate#6 / 2517.03Source ↗official
SafetyPromptsscore#15 / 1598.37Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#31 / 3372.41Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#31 / 3372.64Source ↗official
SALAD-Benchattack_enhanced_malicious_use#31 / 3371.86Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#31 / 3371.38Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#31 / 3375.65Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#31 / 3368.83Source ↗official
SALAD-Benchbase_human_autonomy_integrity#10 / 3389.4Source ↗official
SALAD-Benchbase_information_safety_harms#9 / 3390.79Source ↗official
SALAD-Benchbase_malicious_use#12 / 3388.74Source ↗official
SALAD-Benchbase_misinformation_harms#13 / 3392.61Source ↗official
SALAD-Benchbase_representation_toxicity#13 / 3387.99Source ↗official
SALAD-Benchbase_socioeconomic_harms#6 / 3377.32Source ↗official
SORRY-Benchavg#35 / 510.18Source ↗official
SpeechMap model completioncomplete_pct#6 / 18194.5Source ↗official
SuperCLUE Safetyinstruction_attack#18 / 3168.97Source ↗official
SuperCLUE Safetytraditional_safety#12 / 3173.4Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism30.4
Self-direction23.5
Care / Harm34.3
Fairness / Cheating30.1
Ethical89.2

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)0.479
Completely inaccurate rate7.06