Developer
OpenAI
63 indexed models; 44 currently meet the evidence threshold for the overall ranking. Together they have results from 145 evaluations.
Company governance evidence
OpenAI is represented at +1.10 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by OpenAI
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| GPT 6 Astra | 14 | 6/7 | 1 | 2026-09-03 |
| GPT 5.6 Terra | 25 | 7/7 | 7 | 2026-07-09 |
| GPT 5.3 Chat | 6 | 5/7 | 9 | 2026-03-03 |
| GPT 5.5 Instant | 4 | 3/7 | 12 | 2026-05-26 |
| GPT 5.4 | 37 | 7/7 | 13 | 2026-03-05 |
| GPT 5.5 | 44 | 7/7 | 16 | 2026-04-23 |
| GPT 5.6 Sol | 32 | 7/7 | 21 | 2026-07-09 |
| GPT 5.6 Luna | 27 | 7/7 | 22 | 2026-07-09 |
| GPT 5.2 | 35 | 7/7 | 26 | 2025-12-11 |
| GPT 5.2 Chat | 5 | 4/7 | 27 | 2025-12-11 |
| GPT 5.5 Pro | 3 | 3/7 | 31 | 2026-04-23 |
| GPT 5 Mini | 33 | 7/7 | 33 | 2025-08-07 |
| GPT 5 Nano | 27 | 7/7 | 36 | 2025-08-07 |
| GPT 5 | 44 | 7/7 | 37 | 2025-08-07 |
| GPT 5 Pro | 5 | 4/7 | 39 | 2025-08-07 |
| GPT 5.1 | 34 | 7/7 | 48 | 2025-11-13 |
| GPT 5.1 Codex | 3 | 2/7 | 50 | 2025-11-13 |
| o1 Preview | 4 | 6/7 | 51 | 2024-09-12 |
| GPT 5.4 Pro | 4 | 3/7 | 60 | 2026-03-05 |
| ChatGPT 4o (Mar 2025) | 3 | 3/7 | 62 | 2025-03-27 |
| GPT OSS 20B | 20 | 7/7 | 64 | 2025-08-05 |
| GPT 5.4 Mini | 23 | 6/7 | 69 | 2026-03-17 |
| GPT 4.1 | 37 | 7/7 | 86 | 2025-04-14 |
| o1 Mini | 11 | 7/7 | 94 | 2024-09-12 |
| GPT 4.5 Preview | 11 | 6/7 | 98 | 2025-02-27 |
| o1 | 19 | 6/7 | 100 | 2024-12-17 |
| GPT 5.3 Codex | 5 | 5/7 | 103 | 2026-02-05 |
| o3 | 30 | 6/7 | 109 | 2025-04-16 |
| GPT OSS Safeguard 20B | 4 | 5/7 | 113 | 2025-10-29 |
| GPT 5.2 Codex | 3 | 3/7 | 116 | 2025-12-18 |
| GPT OSS 120B | 35 | 7/7 | 118 | 2025-08-05 |
| GPT 4.1 Mini | 22 | 7/7 | 129 | 2025-04-14 |
| GPT 4 Turbo | 25 | 7/7 | 144 | 2023-11-06 |
| GPT 4o | 73 | 7/7 | 162 | 2024-05-13 |
| Text Davinci 003 | 3 | 4/7 | 169 | 2022-11-28 |
| GPT 4 | 18 | 7/7 | 172 | 2023-03-14 |
| o3 Mini | 26 | 6/7 | 182 | 2025-01-31 |
| GPT 3.5 Turbo | 28 | 7/7 | 185 | 2023-03-01 |
| o4 Mini | 34 | 6/7 | 192 | 2025-04-16 |
| GPT 5.4 Nano | 16 | 6/7 | 201 | 2026-03-17 |
| GPT 4.1 Nano | 15 | 6/7 | 211 | 2025-04-14 |
| GPT 4o Mini | 31 | 7/7 | 257 | 2024-07-18 |
| o3 Pro | 4 | 3/7 | 259 | 2025-06-10 |
| Davinci | 3 | 4/7 | 308 | 2020-06-11 |
| Ada | 1 | 1/7 | — | — |
| Babbage | 1 | 1/7 | — | — |
| ChatGPT | 1 | 1/7 | — | 2022-11-30 |
| Curie | 1 | 1/7 | — | — |
| GPT 3.5 Turbo Instruct | 1 | 1/7 | — | 2023-11-06 |
| GPT 5 Codex | 2 | 1/7 | — | 2025-09-15 |
| GPT 5.1 Chat | 1 | 1/7 | — | 2025-11-13 |
| GPT 5.1 Codex Mini | 1 | 1/7 | — | 2025-11-13 |
| GPT 5.2 Instant | 1 | 3/7 | — | 2025-12-11 |
| GPT 5.2 Pro | 1 | 1/7 | — | 2025-12-10 |
| GPT 5.3 Instant | 1 | 3/7 | — | 2026-03-03 |
| GPT-OSS Safeguard 120B | 2 | 1/7 | — | 2025-09-18 |
| o1 Pro | 2 | 1/7 | — | 2025-03-19 |
| o4-mini Deep Research | 1 | 1/7 | — | 2025-06-26 |
| Text Ada 001 | 1 | 1/7 | — | — |
| Text Babbage 001 | 1 | 1/7 | — | — |
| Text Curie 001 | 1 | 1/7 | — | — |
| Text Davinci 001 | 1 | 4/7 | — | 2022-01-27 |
| Text Davinci 002 | 2 | 4/7 | — | 2022-03-15 |
Evaluations covering OpenAI models (145)
AA-Omniscience · AbstentionBench · Adversarial Humanities Benchmark (AHB) — Table 5 · Adversarial Poetry Refusal (AHB self-run) · Adversarial Poetry — AILuminate Baseline and Poetry ASR · Adversarial Robustness · Agent-SafetyBench · AgentAbstain · AgentDojo · AgentDrive Safety Compliance · AgentHarm · AILuminate General Purpose AI Chat · AIMS Safety-Classifier Competence · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · Arena Factuality — Search Arena (factuality-only weighting) · Arena Factuality — Text Arena (factuality-only weighting) · AuAu Authoritarian Response Audit · BioSecBench-Refusal · BioTIER · BlueBench AttaQ-100 · BrokenMath · BullshitBench v2 · CAIS Risk Index · CASE-Bench · CheatBench direct cheating propensity · Chinese Bias Benchmark for Question Answering · Cisco AI Defense Rolling Single-Turn Leaderboard · Claude Fable 5.1 card — Gray Swan indirect prompt injection k=15 · COMPL-AI AI-Identity Disclosure · COMPL-AI LLM RuLES Multi-Turn Rule Following · COMPL-AI TensorTrust Goal-Hijacking Resistance · Concordia AI Risk Monitor · Confabulations · Constitutional Following — Anthropic Constitution · Constitutional Following — OpenAI Model Spec · Contextual MoralChoice · CRiskEval · CValues · DecodingTrust · DelusionEval · Do-Not-Answer · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · Every Model Cheats — Cybench Cheat Propensity · Fake Alignment (FINE) · FinEval 6.0 Safety Awareness · FinEval Financial Security Knowledge · FlagEval Safety and Values · FLAMES · FORTRESS · Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15 · Governance Decay under Passive Context Compaction · GPT 6 Astra system-card alignment evaluations · GPT-5.6 system card · Gray Swan indirect prompt injection (15 attempts) · HalluVerse-M3 Hallucination Recognition · HarmBench · HarmVideoBench · HELM Classic RealToxicityPrompts · HELM Safety · HUMAINE Trust, Ethics and Safety · Human Pathogen Capabilities Test (HPCT) — overall refusal · Humanity's Last Exam RMS calibration error (Scale Labs) · Inkling-Small model card — FORTRESS · Inkling-Small model card — StrongREJECT · JailBench · JuICE Cultural-Error Span Detection · KIDBench Implicit Child Cue · kindbench v0.1.0 psychological safety ranking · Large-scale Moral Machine experiment on LLMs · LiveSecBench · LLM Ethics Benchmark · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MedSafetyBench · MonitoringBench Full-Trajectory Monitor · MORU · MT-JailBench CrescendoX · MuPPET Contextual Privacy · NESSiE Necessary Safety Benchmark · ODCV-Bench · Olam Social Poker — Social Lie Rate · OpenAgentSafety · OpenAI GPT-4o System Card · OpenAI GPT-5 System Card · OpenAI GPT-5.3 Dynamic Wellbeing · OpenAI GPT-5.4 Dynamic Wellbeing · OpenAI GPT-5.4 First-Person Fairness · OpenAI GPT-5.4 Property Preservation · OpenAI GPT-5.4 User Confirmations · OpenAI o3 and o4-mini System Card · OpenAI o3-mini System Card · Opposite-Narrator Sycophancy · OR-Bench · PacifAIst · PandaBench JBB direct-request panel · Pander Score · PHARE · Pokee-Isaac model card — DTAP · PropensityBench · Qwen2 Safety Panel · RealityTest — Text AI-Identity Disclosure · RefusalBench · Reward Hacking Benchmark · S-Eval · SABER · SafeArena · SafeDialBench · SafetyBench · SafetyPrompts · SALAD-Bench · Shell · SimpleQA Verified · SM-Bench · Social Welfare Function Benchmark · SORRY-Bench · SOSBench · SpeciesismBench · SpeciEval · StereoTales Harmful Associations · SuperCLUE Safety · SYCON Bench · TAC · ThaiSafetyBench · The Dictatorship Eval · ToolPrivacyBench · TrustLLM contemporary collapsed application · TukaBench · UAVBench safety-critical decision recognition · UK AISI cyber-evaluation cheating and prompted self-report · Vals AI Cheating Audit · Vectara HHEM Factual Consistency · VETO Misfired Alignment · Vigil Mental Health Safety · WildClawBench Safety & Alignment (OpenClaw harness) · XSTest
