Developer
Moonshot AI
7 indexed models; 4 currently meet the evidence threshold for the overall ranking. Together they have results from 38 evaluations.
Models by Moonshot AI
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Kimi K3 | 14 | 6/7 | 50 | 2026-07-16 |
| Kimi K2.6 | 20 | 7/7 | 63 | 2026-04-14 |
| Kimi K2 | 30 | 7/7 | 78 | 2025-07-11 |
| Kimi K2.5 | 27 | 7/7 | 128 | 2026-01-01 |
| Kimi Dev 72B | 1 | 1/7 | — | — |
| Kimi K2.7 Code | 2 | 2/7 | — | 2026-06-11 |
| Moonshot V1 | 2 | 3/7 | — | 2024-02-21 |
Evaluations covering Moonshot AI models (38)
AA-Omniscience · AgentAbstain · AgentDrive Safety Compliance · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · Arena Factuality — Text Arena (factuality-only weighting) · BullshitBench v2 · CAIS Risk Index · Cisco AI Defense Rolling Single-Turn Leaderboard · Confabulations · DystopiaBench · Enkrypt AI Safety Leaderboard · FORTRESS · Governance Decay under Passive Context Compaction · HarmVideoBench · HELM Safety · HUMAINE Trust, Ethics and Safety · LiveSecBench · M3-SafetyBench · MACHIAVELLI · Manager Coercion Bench · MASK · ODCV-Bench · Olam Social Poker — Social Lie Rate · PHARE · RealityTest — Text AI-Identity Disclosure · RefusalBench · SABER · SafeDialBench · Shell · SM-Bench · Social Welfare Function Benchmark · SpeciEval · StereoTales Harmful Associations · TAC · ToolPrivacyBench · UAVBench safety-critical decision recognition · Vectara HHEM Factual Consistency
