Model profile
GPT 4.1 Nano
Evidence summary
GPT 4.1 Nano has an estimated overall rank of #174; its 90% source-sensitivity interval is #100–#232. Its behavior-only rank is #187; company governance moves the combined estimate to #174. Published evidence spans 12 evals and 6 of 7 behavior components. Its strongest relative result is HELM Safety (anthropic_red_team, #13 of 80); its weakest is PHARE (harm_resistance_diagnostic, #69 of 70).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗gpt-4-1-nano
- OpenRouter ↗openai/gpt-4.1-nano
- Official model documentation ↗Family-level model document · openai · first party
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Better | Source |
|---|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | #154 / 311 | 0.8041 | ↓ lower | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #49 / 80 | 0.615 | ↑ higher | Source ↗official | |
| BlueBench AttaQ-100attaq_harmlessness_reward_pct | #7 / 18 | 86.67 | ↑ higher | Source ↗official | |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | #78 / 105 | 60.62 | ↓ lower | Source ↗official | |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | #181 / 260 | 12.92 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | #109 / 260 | 89.5 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | #158 / 260 | 60 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | #135 / 258 | 96.14 | ↑ higher | Source ↗official | |
| HELM Safetyanthropic_red_team | #13 / 80 | 0.996 | ↑ higher | Source ↗official | |
| HELM Safetybbq | #60 / 80 | 0.875 | ↑ higher | Source ↗official | |
| HELM Safetyharmbench | #24 / 80 | 0.868 | ↑ higher | Source ↗official | |
| HELM Safetysimple_safety_tests | #32 / 80 | 0.99 | ↑ higher | Source ↗official | |
| HELM Safetyxstest | #39 / 80 | 0.96 | ↑ higher | Source ↗official | |
| MASKlying_probability_pct | #26 / 53 | 38.6 | ↓ lower | Source ↗official | |
| PHAREbias_resistance_diagnostic | #53 / 66 | 0.3622 | ↑ higher | Source ↗official | |
| PHAREhallucination_resistance_diagnostic | #53 / 70 | 0.6774 | ↑ higher | Source ↗official | |
| PHAREharm_resistance_diagnostic | #69 / 70 | 0.7254 | ↑ higher | Source ↗official | |
| PHAREjailbreak_resistance_diagnostic | #29 / 67 | 0.533 | ↑ higher | Source ↗official |
