Developer
DeepSeek
20 indexed models; 13 currently meet the evidence threshold for the overall ranking. Together they have results from 70 evaluations.
Company governance evidence
DeepSeek is represented at -1.45 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by DeepSeek
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Deepseek V4 Flash | 18 | 7/7 | 68 | 2026-04-22 |
| Deepseek V3.2 Exp | 6 | 5/7 | 79 | 2025-09-29 |
| Deepseek V3.2 | 30 | 7/7 | 81 | 2025-12-01 |
| Deepseek V3.1 | 12 | 7/7 | 92 | 2025-08-21 |
| Deepseek V4 Pro | 15 | 7/7 | 114 | 2026-04-22 |
| Deepseek V3.1 Terminus | 3 | 3/7 | 118 | 2025-09-22 |
| Deepseek V3 | 33 | 7/7 | 136 | 2024-12-26 |
| Deepseek R1 | 42 | 7/7 | 175 | 2025-01-20 |
| Deepseek LLM 7B Chat | 3 | 3/7 | 203 | 2023-11-29 |
| DeepSeek V4 Flash 0731 | 4 | 2/7 | 204 | 2026-07-31 |
| Deepseek V3.2 Speciale | 3 | 3/7 | 285 | 2025-11-28 |
| Deepseek LLM 67B Chat | 8 | 5/7 | 288 | 2023-11-29 |
| Deepseek R1 Distill Llama 70B | 4 | 3/7 | 300 | 2025-01-20 |
| Deepseek Chat | 1 | 3/7 | — | — |
| Deepseek R1 0528 Qwen3 8B | 2 | 1/7 | — | 2025-05-29 |
| Deepseek R1 Distill Llama 8B | 1 | 3/7 | — | — |
| Deepseek R1 Distill Qwen 7B | 2 | 3/7 | — | 2025-01-20 |
| Deepseek R1 Zero | 1 | 1/7 | — | — |
| Deepseek V2.5 | 2 | 3/7 | — | 2024-09-05 |
| DeepSeek V4 Pro 0813 | 1 | 1/7 | — | 2026-08-13 |
Evaluations covering DeepSeek models (70)
AA-Omniscience · AbstentionBench · Agent-SafetyBench · AgentAbstain · AgentDrive Safety Compliance · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · Arena Factuality — Text Arena (factuality-only weighting) · AuAu Authoritarian Response Audit · BrokenMath · BullshitBench v2 · CAIS Risk Index · ChineseSafe · ChiSafetyBench · Cisco AI Defense Rolling Single-Turn Leaderboard · Confabulations · Contextual MoralChoice · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · FinEval 6.0 Safety Awareness · FlagEval Safety and Values · FORTRESS · Governance Decay under Passive Context Compaction · HalluVerse-M3 Hallucination Recognition · HELM Safety · HUMAINE Trust, Ethics and Safety · Inkling-Small model card — FORTRESS · Inkling-Small model card — StrongREJECT · JailBench · KIDBench Implicit Child Cue · LiveSecBench · LLM Ethics Benchmark · M3-SafetyBench · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MORU · MT-JailBench CrescendoX · Olam Social Poker — Social Lie Rate · OpenAgentSafety · PandaBench JBB direct-request panel · PHARE · RealityTest — Text AI-Identity Disclosure · RefusalBench · Reward Hacking Benchmark · SABER · SafeDialBench · Shell · SM-Bench · Social Welfare Function Benchmark · SOSBench · SpeciesismBench · SpeciEval · StereoTales Harmful Associations · SYCON Bench · TAC · ToolPrivacyBench · TrustLLM contemporary collapsed application · TukaBench · UAVBench safety-critical decision recognition · Vectara HHEM Factual Consistency · VETO Misfired Alignment · Vigil Mental Health Safety