← Models

Model profile

GPT 4

OpenAIdeveloper
2023-03-14release date
#120 / 309overall rank
15eval lineages

Evidence summary

GPT 4 has an estimated overall rank of #120; its 90% source-sensitivity interval is #41–#214. Its behavior-only rank is #128; company governance moves the combined estimate to #120. Published evidence spans 15 evals and 7 of 7 behavior components. Its strongest relative result is SafetyBench (OFF, #1 of 21); its weakest is DecodingTrust (stereotype_bias, #8 of 8).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#107 / 3280.6941Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#46 / 800.642Source ↗official
DecodingTrustmachine_ethics#4 / 876.6Source ↗official
DecodingTruststereotype_bias#8 / 877Source ↗official
DecodingTrusttoxicity#6 / 841Source ↗official
Do-Not-Answerhuman_harmlessness_rate#4 / 697.55Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#125 / 24115.25Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#58 / 24191.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#82 / 24180.56Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#176 / 23993.09Source ↗official
FLAMESdata_protection#4 / 1350Source ↗official
FLAMESfairness#5 / 1341.37Source ↗official
FLAMESlegality#10 / 1330.43Source ↗official
FLAMESmorality#4 / 1350.75Source ↗official
FLAMESsafety#10 / 1327.51Source ↗official
HarmBenchdr#16 / 2821Source ↗official
JailBenchjailbreak_success_rate#1 / 1434.5Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#3 / 390.7334Source ↗official
MACHIAVELLIdeception_relative_random_pct#29 / 5090Source ↗official
Qwen2 Safety Panelfraud#2 / 36.87Source ↗official
Qwen2 Safety Panelillegal#1 / 30Source ↗official
Qwen2 Safety Panelpornography#2 / 323.63Source ↗official
Qwen2 Safety Panelprivacy#2 / 33.37Source ↗official
SafetyBenchEM#1 / 2192.25Source ↗official
SafetyBenchIA#1 / 2192.35Source ↗official
SafetyBenchMH#1 / 2192.8Source ↗official
SafetyBenchOFF#1 / 2186.15Source ↗official
SafetyBenchPH#1 / 2194.35Source ↗official
SafetyBenchPP#1 / 2191Source ↗official
SafetyBenchUB#1 / 2177.9Source ↗official
SORRY-Benchavg#26 / 510.29Source ↗official
SuperCLUE Safetyinstruction_attack#2 / 3177.59Source ↗official
SuperCLUE Safetyresponsible_ai#3 / 3174.55Source ↗official
SuperCLUE Safetytraditional_safety#23 / 3170.21Source ↗official
XSTestsafe_full_compliance_rate#2 / 30.916Source ↗official
XSTestunsafe_full_refusal_rate#1 / 30.975Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)1.03
Completely inaccurate rate10.7