← Models

Model profile

Grok 4 Fast

xAIdeveloper
2025-09-19release date
#102 / 267overall rank
12eval lineages

Evidence summary

Grok 4 Fast has an estimated overall rank of #102; its 90% source-sensitivity interval is #40–#174. Its behavior-only rank is #100; company governance moves the combined estimate to #102. Published evidence spans 12 evals and 5 of 7 behavior components. Its strongest relative result is PHARE (bias_resistance_diagnostic, #2 of 66); its weakest is CAIS Risk Index (mask, #49 of 51).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#83 / 3110.6596↓ lowerSource ↗official
ANIMAscore#13 / 190.6164↑ higherSource ↗official
BrokenMathsycophancy#4 / 940↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#19 / 4864.1↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#35 / 4968.7↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#41 / 4596.2↓ lowerSource ↗official
CAIS Risk Indexmask#49 / 5170↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#2 / 3230.8↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#29 / 4819.8↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#43 / 10530.03↓ lowerSource ↗official
MACHIAVELLIdeception_relative_random_pct#43 / 5096.2↓ lowerSource ↗official
PHAREbias_resistance_diagnostic#2 / 660.8026↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#62 / 700.634↑ higherSource ↗official
PHAREharm_resistance_diagnostic#64 / 700.8134↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#49 / 670.412↑ higherSource ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#14 / 270.695↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-11.2
Government46
Diplomacy62.5
Economy40
Society63.8