← Models

Model profile

Glm 4.5

Z.aideveloper
2025-07-28release date
#199 / 309overall rank
9eval lineages

Evidence summary

Glm 4.5 has an estimated overall rank of #199; its 90% source-sensitivity interval is #63–#263. Its behavior-only rank is #190; company governance moves the combined estimate to #199. Published evidence spans 9 evals and 7 of 7 behavior components. Its strongest relative result is Confabulations (confabulation_rate, #8 of 52); its weakest is FlagEval Safety and Values (a1_qualified_rate, #18 of 18).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#110 / 3280.7015Source ↗official
BullshitBench v2clear_pushback_rate#94 / 1060.07Source ↗official
Confabulationsconfabulation_rate#8 / 527.921Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#46 / 24130.49Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#209 / 24177.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#172 / 24152.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#134 / 23995.76Source ↗official
FlagEval Safety and Valuesa1_qualified_rate#18 / 1859.1Source ↗official
FlagEval Safety and Valuesa2_qualified_rate#18 / 1857.84Source ↗official
FlagEval Safety and Valuesa3_qualified_rate#18 / 1860.32Source ↗official
FlagEval Safety and Valuesa4_qualified_rate#18 / 1860.7Source ↗official
FlagEval Safety and Valuesa5_qualified_rate#17 / 1856.79Source ↗official
FORTRESSaverage_risk_score#44 / 4959.58Source ↗official
FORTRESSover_refusal_score#15 / 462.54Source ↗official
MASKlying_probability_pct#25 / 5338.54Source ↗official
Social Welfare Function Benchmarkfairness#13 / 190.475Source ↗official
SpeciEvalbelief_animal_sentience#69 / 1136.7Source ↗official
SpeciEvalland_animal_4ns#49 / 1134.47Source ↗official
SpeciEvalsea_animal_4ns#98 / 1135.09Source ↗official
SpeciEvalspeciesism#23 / 1131.58Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-18.1
Government46.5
Diplomacy66.2
Economy45.2
Society60.8

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions2.81
CCP-narrative alignment — China topics3.34
CCP-narrative alignment — non-China controls1.22