Developer
Z.ai
33 indexed models; 15 currently meet the evidence threshold for the overall ranking. Together they have results from 72 evaluations.
Company governance evidence
Z.ai is represented at -0.06 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by Z.ai
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| GLM 5.3 | 14 | 6/7 | 25 | 2026-08-18 |
| GLM 4.5 Air | 14 | 5/7 | 68 | 2025-07-28 |
| GLM 5.3 Flash | 5 | 6/7 | 71 | 2026-08-26 |
| GLM 5 | 16 | 7/7 | 82 | 2026-02-11 |
| GLM 5.2 | 25 | 7/7 | 84 | 2026-06-16 |
| GLM 5.1 | 20 | 7/7 | 126 | 2026-04-03 |
| GLM 4.7 Flash | 6 | 3/7 | 174 | 2026-01-19 |
| GLM 4 9B Chat | 8 | 4/7 | 196 | 2024-06-05 |
| GLM 4.7 | 9 | 6/7 | 220 | 2026-06-03 |
| GLM 4.6 | 9 | 6/7 | 223 | 2025-09-30 |
| GLM 4.5 | 11 | 7/7 | 233 | 2025-07-28 |
| GLM 5 Turbo | 7 | 6/7 | 237 | 2026-03-15 |
| ChatGLM 6B | 5 | 4/7 | 244 | 2023-03-14 |
| ChatGLM3 6B | 10 | 4/7 | 262 | 2023-10-27 |
| ChatGLM2 6B | 6 | 3/7 | 287 | 2023-06-25 |
| ChatGLM 130B | 1 | 4/7 | — | 2023-03-14 |
| ChatGLM Turbo | 1 | 3/7 | — | — |
| ChatGLM2 Lite | 1 | 3/7 | — | 2023-06-25 |
| ChatGLM3 | 1 | 1/7 | — | 2023-10-27 |
| ChatGLM3 6B 32K | 1 | 1/7 | — | 2023-10-27 |
| ChatGLM3 Turbo | 1 | 3/7 | — | 2023-10-27 |
| ChatGLM4 | 1 | 3/7 | — | 2024-01-16 |
| GLM | 1 | 1/7 | — | — |
| GLM 10B | 1 | 1/7 | — | 2023-02-28 |
| GLM 130B | 1 | 1/7 | — | — |
| GLM 350M | 1 | 1/7 | — | — |
| GLM 4 Flash | 1 | 1/7 | — | — |
| GLM 4 Plus | 1 | 1/7 | — | — |
| GLM 4.5V | 2 | 2/7 | — | 2025-08-10 |
| GLM 4.6V | 2 | 1/7 | — | 2025-12-07 |
| GLM 5V Turbo | 2 | 1/7 | — | 2026-04-01 |
| GLM Z1 32B | 1 | 1/7 | — | 2025-04-08 |
| LongWriter-GLM4 9B | 1 | 3/7 | — | 2024-08-12 |
Evaluations covering Z.ai models (72)
AA-Omniscience · Adversarial Humanities Benchmark (AHB) — Table 5 · Adversarial Poetry Refusal (AHB self-run) · Agent-SafetyBench · AgentAbstain · AgentDrive Safety Compliance · AIRBench 2024 Safety Scenarios · ANIMA · Arena Factuality — Text Arena (factuality-only weighting) · BioTIER · BullshitBench v2 · CAIS Risk Index · Chinese Bias Benchmark for Question Answering · ChineseSafe · ChiSafetyBench · Cisco AI Defense Rolling Single-Turn Leaderboard · CMoralEval · Concordia AI Risk Monitor · Confabulations · CRiskEval · CValues · Do-Not-Answer · DSPSafeBench · DystopiaBench · Enkrypt AI Safety Leaderboard · Every Model Cheats — Cybench Cheat Propensity · Fake Alignment (FINE) · FinEval 6.0 Safety Awareness · FinEval Financial Security Knowledge · FlagEval Safety and Values · FLAMES · FORTRESS · Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15 · Governance Decay under Passive Context Compaction · HarmVideoBench · HELM Classic RealToxicityPrompts · HELM Safety · HUMAINE Trust, Ethics and Safety · Human Pathogen Capabilities Test (HPCT) — overall refusal · Humanity's Last Exam RMS calibration error (Scale Labs) · JailBench · LiveSecBench · M3-SafetyBench · MACHIAVELLI · Manager Coercion Bench · MASK · ODCV-Bench · Olam Social Poker — Social Lie Rate · Opposite-Narrator Sycophancy · PandaBench JBB direct-request panel · Pander Score · PHARE · RefusalBench · S-Eval · SABER · SafeDialBench · SafetyBench · SafetyPrompts · SALAD-Bench · SM-Bench · Social Welfare Function Benchmark · SORRY-Bench · SpeciEval · StereoTales Harmful Associations · SuperCLUE Safety · TAC · The Dictatorship Eval · ToolPrivacyBench · TrustLLM contemporary collapsed application · Vals AI Cheating Audit · Vectara HHEM Factual Consistency · WildClawBench Safety & Alignment (OpenClaw harness)