Developer
Moonshot AI
6 indexed models; 4 currently meet the evidence threshold for the overall ranking. Together they have results from 30 evaluations.
Models by Moonshot AI
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Kimi K3 | 13 | 6/7 | 45 | 2026-07-16 |
| Kimi K2.6 | 18 | 7/7 | 63 | 2026-04-14 |
| Kimi K2 | 26 | 7/7 | 74 | 2025-07-11 |
| Kimi K2.5 | 22 | 7/7 | 85 | 2026-01-01 |
| Kimi K2.7 Code | 2 | 2/7 | — | 2026-06-11 |
| Moonshot V1 | 2 | 3/7 | — | 2024-02-21 |
Evaluations covering Moonshot AI models (30)
AA-Omniscience · AgentAbstain · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · BullshitBench v2 · CAIS Risk Index · Cisco AI Defense Rolling Single-Turn Leaderboard · Confabulations · DystopiaBench · Enkrypt AI Safety Leaderboard · FORTRESS · HELM Safety · HUMAINE Trust, Ethics and Safety · LiveSecBench · M3-SafetyBench · MACHIAVELLI · Manager Coercion Bench · MASK · ODCV-Bench · PHARE · RefusalBench · SABER · SafeDialBench · Shell · SM-Bench · Social Welfare Function Benchmark · SpeciEval · TAC · ToolPrivacyBench · UAVBench safety-critical decision recognition
