Developer
OpenAI
62 indexed models; 41 currently meet the evidence threshold for the overall ranking. Together they have results from 122 evaluations.
Company governance evidence
OpenAI is represented at +1.09 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by OpenAI
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| GPT 5.6 Terra | 19 | 7/7 | 4 | 2026-07-09 |
| GPT-5.5 Instant | 4 | 3/7 | 5 | 2026-05-26 |
| GPT-5.3 Chat | 6 | 5/7 | 9 | 2026-03-03 |
| GPT 5.4 | 27 | 7/7 | 13 | 2026-03-05 |
| GPT 5 Nano | 23 | 7/7 | 14 | 2025-08-07 |
| GPT 5.5 | 36 | 7/7 | 16 | 2026-04-23 |
| GPT 5.2 Chat | 5 | 4/7 | 19 | 2025-12-11 |
| GPT 5.2 | 31 | 7/7 | 23 | 2025-12-11 |
| GPT 5 Mini | 29 | 7/7 | 25 | 2025-08-07 |
| GPT 5.6 Luna | 21 | 7/7 | 26 | 2026-07-09 |
| GPT 5.6 Sol | 20 | 7/7 | 28 | 2026-07-09 |
| GPT 5.1 Codex | 3 | 2/7 | 29 | — |
| O1 Preview | 4 | 6/7 | 41 | 2024-09-12 |
| ChatGPT-4o | 3 | 3/7 | 42 | 2025-03-27 |
| GPT 5.1 | 30 | 7/7 | 44 | 2025-11-13 |
| GPT 4.5 Preview | 9 | 6/7 | 48 | 2025-02-27 |
| GPT 5 | 38 | 7/7 | 49 | 2025-08-07 |
| GPT 5 Pro | 3 | 4/7 | 64 | 2025-08-07 |
| GPT 4.1 | 31 | 7/7 | 72 | 2025-04-14 |
| GPT 5.2 Codex | 3 | 3/7 | 76 | — |
| GPT 5.3 Codex | 5 | 5/7 | 85 | 2026-02-05 |
| GPT Oss Safeguard 20B | 4 | 5/7 | 91 | 2025-10-29 |
| O1 Mini | 11 | 7/7 | 95 | 2024-09-12 |
| O3 | 26 | 6/7 | 96 | 2025-04-16 |
| GPT Oss 20B | 16 | 7/7 | 103 | 2025-08-05 |
| O1 | 17 | 6/7 | 108 | 2024-12-17 |
| GPT Oss 120B | 31 | 7/7 | 113 | 2025-08-05 |
| GPT 5.4 Mini | 18 | 6/7 | 116 | 2026-03-17 |
| GPT 4 | 15 | 7/7 | 120 | 2023-03-14 |
| GPT 4.1 Mini | 19 | 7/7 | 121 | 2025-04-14 |
| GPT 5.4 Pro | 3 | 3/7 | 123 | — |
| GPT 4 Turbo | 21 | 7/7 | 126 | 2023-11-06 |
| GPT 4O | 67 | 7/7 | 144 | 2024-05-13 |
| GPT 3.5 Turbo | 24 | 7/7 | 150 | 2023-03-01 |
| Text Davinci 003 | 3 | 4/7 | 151 | 2022-11-28 |
| O3 Mini | 24 | 6/7 | 188 | 2025-01-31 |
| O4 Mini | 29 | 6/7 | 189 | 2025-04-16 |
| GPT 4.1 Nano | 13 | 6/7 | 192 | 2025-04-14 |
| GPT 5.4 Nano | 14 | 5/7 | 198 | 2026-03-17 |
| GPT 4O Mini | 28 | 7/7 | 237 | 2024-07-18 |
| Davinci | 3 | 4/7 | 284 | 2020-06-11 |
| Ada | 1 | 1/7 | — | — |
| Babbage | 1 | 1/7 | — | — |
| Chatgpt | 1 | 1/7 | — | 2022-11-30 |
| Curie | 1 | 1/7 | — | — |
| GPT 3.5 Turbo Instruct | 1 | 1/7 | — | — |
| GPT 5 Codex | 2 | 1/7 | — | — |
| GPT 5.1 Codex Mini | 1 | 1/7 | — | — |
| GPT 5.2 Instant | 1 | 3/7 | — | — |
| GPT 5.3 Instant | 1 | 3/7 | — | — |
| GPT-5.1 Chat | 1 | 1/7 | — | 2025-11-13 |
| GPT-5.2 Pro | 1 | 1/7 | — | 2025-12-10 |
| GPT-5.5 Pro | 1 | 1/7 | — | 2026-04-23 |
| GPT-OSS Safeguard 120B | 2 | 1/7 | — | — |
| O1 Pro | 1 | 1/7 | — | — |
| O3 Pro | 3 | 1/7 | — | — |
| o4-mini Deep Research | 1 | 1/7 | — | 2025-06-26 |
| Text Ada 001 | 1 | 1/7 | — | — |
| Text Babbage 001 | 1 | 1/7 | — | — |
| Text Curie 001 | 1 | 1/7 | — | — |
| Text Davinci 001 | 1 | 4/7 | — | 2022-01-27 |
| Text Davinci 002 | 2 | 4/7 | — | 2022-03-15 |
Evaluations covering OpenAI models (122)
AA-Omniscience · AbstentionBench · Adversarial Robustness · Agent-SafetyBench · AgentAbstain · AgentDojo · AgentDrive Safety Compliance · AgentHarm · AILuminate General Purpose AI Chat · AIMS Safety-Classifier Competence · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · Arena Factuality — Search Arena (factuality-only weighting) · Arena Factuality — Text Arena (factuality-only weighting) · AuAu Authoritarian Response Audit · BioSecBench-Refusal · BlueBench AttaQ-100 · BrokenMath · BullshitBench v2 · CAIS Risk Index · CASE-Bench · Chinese Bias Benchmark for Question Answering · Cisco AI Defense Rolling Single-Turn Leaderboard · COMPL-AI AI-Identity Disclosure · COMPL-AI LLM RuLES Multi-Turn Rule Following · COMPL-AI TensorTrust Goal-Hijacking Resistance · Confabulations · Constitutional Following — Anthropic Constitution · Constitutional Following — OpenAI Model Spec · Contextual MoralChoice · CRiskEval · CValues · DecodingTrust · Do-Not-Answer · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · Fake Alignment (FINE) · FinEval 6.0 Safety Awareness · FinEval Financial Security Knowledge · FlagEval Safety and Values · FLAMES · FORTRESS · Governance Decay under Passive Context Compaction · GPT-5.6 system card · Gray Swan indirect prompt injection (15 attempts) · HalluVerse-M3 Hallucination Recognition · HarmBench · HarmVideoBench · HELM Classic RealToxicityPrompts · HELM Safety · HUMAINE Trust, Ethics and Safety · Inkling-Small model card — FORTRESS · Inkling-Small model card — StrongREJECT · JailBench · JuICE Cultural-Error Span Detection · KIDBench Implicit Child Cue · Large-scale Moral Machine experiment on LLMs · LiveSecBench · LLM Ethics Benchmark · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MonitoringBench Full-Trajectory Monitor · MORU · MT-JailBench CrescendoX · MuPPET Contextual Privacy · ODCV-Bench · Olam Social Poker — Social Lie Rate · OpenAgentSafety · OpenAI GPT-4o System Card · OpenAI GPT-5 System Card · OpenAI GPT-5.3 Dynamic Wellbeing · OpenAI GPT-5.4 Dynamic Wellbeing · OpenAI GPT-5.4 First-Person Fairness · OpenAI GPT-5.4 Property Preservation · OpenAI GPT-5.4 User Confirmations · OpenAI o3 and o4-mini System Card · OpenAI o3-mini System Card · OR-Bench · PacifAIst · PandaBench JBB direct-request panel · PHARE · Pokee-Isaac model card — DTAP · PropensityBench · Qwen2 Safety Panel · RealityTest — Text AI-Identity Disclosure · RefusalBench · Reward Hacking Benchmark · S-Eval · SABER · SafeArena · SafeDialBench · SafetyBench · SafetyPrompts · SALAD-Bench · Shell · SM-Bench · Social Welfare Function Benchmark · SORRY-Bench · SOSBench · SpeciesismBench · SpeciEval · StereoTales Harmful Associations · SuperCLUE Safety · SYCON Bench · TAC · ThaiSafetyBench · ToolPrivacyBench · TrustLLM contemporary collapsed application · TukaBench · UAVBench safety-critical decision recognition · Vectara HHEM Factual Consistency · VETO Misfired Alignment · Vigil Mental Health Safety · XSTest
