Developer
xAI
15 indexed models; 10 currently meet the evidence threshold for the overall ranking. Together they have results from 42 evaluations.
Company governance evidence
xAI is represented at -1.25 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by xAI
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Grok 4.20 Multi Agent | 4 | 4/7 | 29 | 2026-03-10 |
| Grok 4.5 | 13 | 6/7 | 65 | 2026-07-08 |
| Grok 4.20 | 18 | 7/7 | 68 | 2026-03-10 |
| Grok 3 Mini | 15 | 7/7 | 87 | 2025-04-03 |
| Grok 4.1 Fast | 16 | 7/7 | 95 | 2025-11-19 |
| Grok 4.3 | 19 | 7/7 | 98 | 2026-05-15 |
| Grok 4 Fast | 12 | 5/7 | 102 | 2025-09-19 |
| Grok 4 | 27 | 7/7 | 105 | 2025-07-09 |
| Grok 3 | 17 | 7/7 | 163 | 2025-04-03 |
| Grok 2 | 4 | 3/7 | 259 | 2024-12-12 |
| Grok 3 Beta | 1 | 2/7 | — | — |
| Grok 3 Eai Hardened System Prompt | 1 | 3/7 | — | — |
| Grok 4 1 Fast Non Reasoning | 2 | 3/7 | — | — |
| Grok Build 0.1 | 2 | 2/7 | — | 2026-05-20 |
| Grok Code Fast 1 | 2 | 2/7 | — | — |
Evaluations covering xAI models (42)
AA-Omniscience · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · BioSecBench-Refusal · BrokenMath · BullshitBench v2 · CAIS Risk Index · Cisco AI Defense Rolling Single-Turn Leaderboard · Confabulations · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · FlagEval Safety and Values · Gray Swan indirect prompt injection (15 attempts) · HELM Safety · HUMAINE Trust, Ethics and Safety · LiveSecBench · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MORU · ODCV-Bench · PacifAIst · PandaBench JBB direct-request panel · PHARE · RefusalBench · Shell · SM-Bench · Social Welfare Function Benchmark · SOSBench · SpeciesismBench · SpeciEval · TrustLLM contemporary collapsed application · TukaBench · UAVBench safety-critical decision recognition · VETO Misfired Alignment · Vigil Mental Health Safety