Developer
DeepSeek
22 indexed models; 14 currently meet the evidence threshold for the overall ranking. Together they have results from 82 evaluations.
Company governance evidence
DeepSeek is represented at -1.44 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by DeepSeek
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | 5 | 3/7 | 28 | 2026-09-10 |
| DeepSeek v4 Flash | 20 | 7/7 | 91 | 2026-04-22 |
| DeepSeek v3.1 | 15 | 7/7 | 106 | 2025-08-21 |
| DeepSeek v3.2 | 34 | 7/7 | 124 | 2025-12-01 |
| DeepSeek v3.2 Exp | 7 | 5/7 | 125 | 2025-09-29 |
| DeepSeek v4 Pro | 23 | 7/7 | 141 | 2026-04-22 |
| DeepSeek v3 | 34 | 7/7 | 150 | 2024-12-26 |
| DeepSeek V4 Flash 0731 | 5 | 3/7 | 177 | 2026-07-31 |
| DeepSeek R1 | 48 | 7/7 | 191 | 2025-01-20 |
| DeepSeek LLM 7B Chat | 3 | 3/7 | 218 | 2023-11-29 |
| DeepSeek v3.1 Terminus | 4 | 4/7 | 230 | 2025-09-22 |
| DeepSeek v3.2 Speciale | 3 | 3/7 | 305 | 2025-11-28 |
| DeepSeek LLM 67B Chat | 8 | 5/7 | 310 | 2023-11-29 |
| DeepSeek R1 Distill Llama 70B | 4 | 3/7 | 325 | 2025-01-20 |
| DeepSeek Chat | 1 | 3/7 | — | — |
| DeepSeek R1 Distill Llama 8B | 1 | 3/7 | — | 2025-01-20 |
| DeepSeek R1 Distill Qwen 7B | 2 | 3/7 | — | 2025-01-20 |
| DeepSeek R1 Qwen3 8B | 2 | 1/7 | — | 2025-05-29 |
| DeepSeek R1 Zero | 1 | 1/7 | — | 2025-01-20 |
| DeepSeek v2.5 | 2 | 3/7 | — | 2024-09-05 |
| DeepSeek V4 Flash Vision | 1 | 1/7 | — | 2026-08-21 |
| DeepSeek V4 Pro 0813 | 2 | 1/7 | — | 2026-08-13 |
Evaluations covering DeepSeek models (82)
AA-Omniscience · AbstentionBench · Adversarial Humanities Benchmark (AHB) — Table 5 · Adversarial Poetry — AILuminate Baseline and Poetry ASR · Agent-SafetyBench · AgentAbstain · AgentDrive Safety Compliance · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · Arena Factuality — Text Arena (factuality-only weighting) · AuAu Authoritarian Response Audit · BioTIER · BrokenMath · BullshitBench v2 · CAIS Risk Index · ChineseSafe · ChiSafetyBench · Cisco AI Defense Rolling Single-Turn Leaderboard · Concordia AI Risk Monitor · Confabulations · Contextual MoralChoice · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · Every Model Cheats — Cybench Cheat Propensity · FinEval 6.0 Safety Awareness · FlagEval Safety and Values · FORTRESS · Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15 · Governance Decay under Passive Context Compaction · HalluVerse-M3 Hallucination Recognition · HELM Safety · HUMAINE Trust, Ethics and Safety · Human Pathogen Capabilities Test (HPCT) — overall refusal · Inkling-Small model card — FORTRESS · Inkling-Small model card — StrongREJECT · JailBench · KIDBench Implicit Child Cue · LiveSecBench · LLM Ethics Benchmark · M3-SafetyBench · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MORU · MT-JailBench CrescendoX · Olam Social Poker — Social Lie Rate · OpenAgentSafety · Opposite-Narrator Sycophancy · PandaBench JBB direct-request panel · PHARE · RealityTest — Text AI-Identity Disclosure · RefusalBench · Reward Hacking Benchmark · SABER · SafeDialBench · Shell · SimpleQA Verified · SM-Bench · Social Welfare Function Benchmark · SOSBench · SpeciesismBench · SpeciEval · StereoTales Harmful Associations · SYCON Bench · TAC · The Dictatorship Eval · ToolPrivacyBench · TrustLLM contemporary collapsed application · TukaBench · UAVBench safety-critical decision recognition · Vals AI Cheating Audit · Vectara HHEM Factual Consistency · VETO Misfired Alignment · Vigil Mental Health Safety · WildClawBench Safety & Alignment (OpenClaw harness)