Developer
xAI
14 indexed models; 11 currently meet the evidence threshold for the overall ranking. Together they have results from 52 evaluations.
Company governance evidence
xAI is represented at -1.25 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by xAI
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Grok 4.20 Multi Agent | 4 | 4/7 | 37 | 2026-03-10 |
| Grok 4.5 | 15 | 6/7 | 58 | 2026-07-08 |
| Grok 4.20 | 19 | 7/7 | 75 | 2026-03-10 |
| Grok 3 Mini | 16 | 7/7 | 99 | 2025-04-03 |
| Grok 4.3 | 20 | 7/7 | 104 | 2026-05-15 |
| Grok 4 | 30 | 7/7 | 138 | 2025-07-09 |
| Grok 4.6 | 3 | 4/7 | 146 | 2026-08-12 |
| Grok 3 | 18 | 7/7 | 181 | 2025-04-03 |
| Grok 4 Fast | 15 | 5/7 | 184 | 2025-09-19 |
| Grok 4.1 Fast | 21 | 7/7 | 206 | 2025-11-19 |
| Grok 2 | 4 | 3/7 | 302 | 2024-12-12 |
| Grok 3 Beta | 1 | 2/7 | — | — |
| Grok Build 0.1 | 2 | 2/7 | — | 2026-05-20 |
| Grok Code Fast 1 | 2 | 2/7 | — | — |
Evaluations covering xAI models (52)
AA-Omniscience · AgentDrive Safety Compliance · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · Arena Factuality — Search Arena (factuality-only weighting) · Arena Factuality — Text Arena (factuality-only weighting) · AuAu Authoritarian Response Audit · BioSecBench-Refusal · BrokenMath · BullshitBench v2 · CAIS Risk Index · Cisco AI Defense Rolling Single-Turn Leaderboard · Confabulations · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · FlagEval Safety and Values · Gray Swan indirect prompt injection (15 attempts) · HELM Safety · HUMAINE Trust, Ethics and Safety · LiveSecBench · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MORU · MT-JailBench CrescendoX · ODCV-Bench · Olam Social Poker — Social Lie Rate · PacifAIst · PandaBench JBB direct-request panel · PHARE · RealityTest — Text AI-Identity Disclosure · RefusalBench · Shell · SM-Bench · Social Welfare Function Benchmark · SOSBench · SpeciesismBench · SpeciEval · StereoTales Harmful Associations · TAC · TrustLLM contemporary collapsed application · TukaBench · UAVBench safety-critical decision recognition · Vectara HHEM Factual Consistency · VETO Misfired Alignment · Vigil Mental Health Safety