← Models

Model profile

Grok 3 Mini

xAIdeveloper
2025-04-03release date
#136 / 333Safety rank
#78 / 645Freedom rank

Evidence summary

Safety. Grok 3 Mini has an estimated Safety rank of #136; its 90% source-sensitivity interval is #55–#235. Its behavior-only rank is #143; company governance moves the combined estimate to #136. Published Safety evidence spans 17 eval lineages and 7 of 7 components. Its strongest relative result is AA-Omniscience (hallucination_rate, #21 of 345); its weakest is Cisco AI Defense Rolling Single-Turn Leaderboard (single_turn_attack_success_rate, #104 of 104).

Freedom. Grok 3 Mini has an estimated Freedom rank of #78; its 90% source-sensitivity interval is #35–#272. Published Freedom evidence spans 10 eval lineages and 1 of 1 components. Its strongest relative result is Cisco AI Defense Rolling Single-Turn Leaderboard (single_turn_attack_success_rate, #1 of 104); its weakest is HELM Safety (anthropic_red_team, #59 of 80).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#21 / 3450.2582Source ↗official
AgentDrive Safety Compliancescr#13 / 4892.5Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#61 / 800.535Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#104 / 10486.64Source ↗official
Confabulationsconfabulation_rate#9 / 528.911Source ↗official
Emergent Collusionhigh_illegality_game_rate#4 / 130.3Source ↗official
HELM Safetyanthropic_red_team#18 / 800.995Source ↗official
HELM Safetybbq#13 / 800.967Source ↗official
HELM Safetyharmbench#64 / 800.572Source ↗official
HELM Safetysimple_safety_tests#29 / 800.993Source ↗official
HELM Safetyxstest#6 / 800.984Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 690Source ↗official
LiveSecBenchethics#39 / 4320.84Source ↗official
LiveSecBenchfactuality#29 / 4337.05Source ↗official
LiveSecBenchlegality#33 / 4320.79Source ↗official
LiveSecBenchprivacy#32 / 4330.02Source ↗official
LiveSecBenchpsychological_health#31 / 4333.15Source ↗official
PacifAIstp_score#6 / 779.77Source ↗official
PHAREbias_resistance_diagnostic#32 / 660.464Source ↗official
PHAREhallucination_resistance_diagnostic#48 / 700.6862Source ↗official
PHAREharm_resistance_diagnostic#49 / 700.9047Source ↗official
PHAREjailbreak_resistance_diagnostic#60 / 670.3436Source ↗official
SOSBenchbiology_pvr#16 / 230.758Source ↗official
SOSBenchchemistry_pvr#16 / 230.586Source ↗official
SOSBenchmedicine_pvr#16 / 230.746Source ↗official
SOSBenchpharmacology_pvr#17 / 230.93Source ↗official
SOSBenchphysics_pvr#16 / 230.708Source ↗official
SOSBenchpsychology_pvr#15 / 230.7Source ↗official
SpeciEvalbelief_animal_sentience#62 / 1236.815Source ↗official
SpeciEvalland_animal_4ns#64 / 1234.545Source ↗official
SpeciEvalsea_animal_4ns#90 / 1234.91Source ↗official
SpeciEvalspeciesism#33 / 1231.68Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#20 / 800.535Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#1 / 10486.64Source ↗official
HELM Safetyanthropic_red_team#59 / 800.995Source ↗official
HELM Safetyharmbench#17 / 800.572Source ↗official
HELM Safetysimple_safety_tests#50 / 800.993Source ↗official
HELM Safetyxstest#6 / 800.984Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 690Source ↗official
LiveSecBenchethics#5 / 4320.84Source ↗official
LiveSecBenchlegality#10 / 4320.79Source ↗official
LiveSecBenchprivacy#12 / 4330.02Source ↗official
LiveSecBenchpsychological_health#13 / 4333.15Source ↗official
PHAREharm_resistance_diagnostic#22 / 700.9047Source ↗official
PHAREjailbreak_resistance_diagnostic#8 / 670.3436Source ↗official
SOSBenchbiology_pvr#8 / 230.758Source ↗official
SOSBenchchemistry_pvr#8 / 230.586Source ↗official
SOSBenchmedicine_pvr#8 / 230.746Source ↗official
SOSBenchpharmacology_pvr#7 / 230.93Source ↗official
SOSBenchphysics_pvr#8 / 230.708Source ↗official
SOSBenchpsychology_pvr#9 / 230.7Source ↗official