Developer
Z.ai
30 indexed models; 13 currently meet the evidence threshold for the overall ranking. Together they have results from 59 evaluations.
Company governance evidence
Z.ai is represented at -0.07 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by Z.ai
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Glm 4.5 Air | 13 | 5/7 | 51 | 2025-07-28 |
| Glm 5.2 | 17 | 6/7 | 93 | 2026-06-16 |
| Glm 5 | 13 | 7/7 | 100 | 2026-02-11 |
| Glm 5.1 | 18 | 6/7 | 105 | 2026-04-03 |
| Glm 4.6 | 8 | 6/7 | 161 | 2025-09-30 |
| Glm 4.7 Flash | 6 | 3/7 | 169 | 2026-01-19 |
| Glm 4 9B Chat | 8 | 4/7 | 179 | 2024-06-05 |
| Glm 4.5 | 9 | 7/7 | 199 | 2025-07-28 |
| Glm 4.7 | 8 | 5/7 | 209 | 2026-06-03 |
| Glm 5 Turbo | 4 | 5/7 | 210 | 2026-03-15 |
| Chatglm 6B | 5 | 4/7 | 218 | 2023-03-14 |
| Chatglm3 6B | 10 | 4/7 | 232 | 2023-10-27 |
| Chatglm2 6B | 6 | 3/7 | 260 | 2023-06-25 |
| Chatglm 130B | 1 | 4/7 | — | 2023-03-14 |
| Chatglm Turbo | 1 | 3/7 | — | — |
| Chatglm2 Lite | 1 | 3/7 | — | 2023-06-25 |
| Chatglm3 | 1 | 1/7 | — | 2023-10-27 |
| Chatglm3 6B 32K | 1 | 1/7 | — | 2023-10-27 |
| Chatglm3 Turbo | 1 | 3/7 | — | 2023-10-27 |
| Chatglm4 | 1 | 3/7 | — | 2024-01-16 |
| Glm | 1 | 1/7 | — | — |
| Glm 10B | 1 | 1/7 | — | — |
| Glm 130B | 1 | 1/7 | — | — |
| Glm 350M | 1 | 1/7 | — | — |
| Glm 4 Flash | 1 | 1/7 | — | — |
| Glm 4 Plus | 1 | 1/7 | — | — |
| Glm 4.5V | 2 | 2/7 | — | 2025-08-10 |
| Glm 4.6V | 2 | 1/7 | — | 2025-12-07 |
| Glm 5V Turbo | 2 | 1/7 | — | — |
| Glm Z1 32B 0414 | 1 | 1/7 | — | — |
Evaluations covering Z.ai models (59)
AA-Omniscience · Agent-SafetyBench · AgentAbstain · AgentDrive Safety Compliance · AIRBench 2024 Safety Scenarios · ANIMA · Arena Factuality — Text Arena (factuality-only weighting) · BullshitBench v2 · CAIS Risk Index · Chinese Bias Benchmark for Question Answering · ChineseSafe · ChiSafetyBench · Cisco AI Defense Rolling Single-Turn Leaderboard · CMoralEval · Confabulations · CRiskEval · CValues · Do-Not-Answer · DSPSafeBench · DystopiaBench · Enkrypt AI Safety Leaderboard · Fake Alignment (FINE) · FinEval 6.0 Safety Awareness · FinEval Financial Security Knowledge · FlagEval Safety and Values · FLAMES · FORTRESS · Governance Decay under Passive Context Compaction · HarmVideoBench · HELM Classic RealToxicityPrompts · HELM Safety · HUMAINE Trust, Ethics and Safety · JailBench · LiveSecBench · M3-SafetyBench · MACHIAVELLI · Manager Coercion Bench · MASK · ODCV-Bench · Olam Social Poker — Social Lie Rate · PandaBench JBB direct-request panel · PHARE · RefusalBench · S-Eval · SABER · SafeDialBench · SafetyBench · SafetyPrompts · SALAD-Bench · SM-Bench · Social Welfare Function Benchmark · SORRY-Bench · SpeciEval · StereoTales Harmful Associations · SuperCLUE Safety · TAC · ToolPrivacyBench · TrustLLM contemporary collapsed application · Vectara HHEM Factual Consistency