← Models

Model profile

OLMo 2 32B Instruct

Ai2developer
2025-03-13release date
#246 / 333Safety rank
#452 / 645Freedom rank

Evidence summary

Safety. OLMo 2 32B Instruct has an estimated Safety rank of #246; its 90% source-sensitivity interval is #81–#288. Its behavior-only rank is #249; company governance moves the combined estimate to #246. Published Safety evidence spans 8 eval lineages and 6 of 7 components. Its strongest relative result is AILuminate General Purpose AI Chat (child_sexual_exploitation, #6 of 32); its weakest is HELM Safety (bbq, #74 of 80).

Freedom. OLMo 2 32B Instruct has an estimated Freedom rank of #452; its 90% source-sensitivity interval is #292–#547. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is HELM Safety (simple_safety_tests, #31 of 80); its weakest is AILuminate General Purpose AI Chat (child_sexual_exploitation, #27 of 32).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AgentDrive Safety Compliancescr#29 / 4876.25Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#6 / 329.1Source ↗official
AILuminate General Purpose AI Chatdefamation#10 / 3210.6Source ↗official
AILuminate General Purpose AI Chathate#8 / 324Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#11 / 3218.8Source ↗official
AILuminate General Purpose AI Chatintellectual_property#11 / 328.2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#9 / 329Source ↗official
AILuminate General Purpose AI Chatprivacy#10 / 328.1Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#11 / 329.2Source ↗official
AILuminate General Purpose AI Chatsexual_content#15 / 3221Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#12 / 3213.9Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#10 / 329.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#11 / 3212.7Source ↗official
HELM Safetyanthropic_red_team#27 / 800.993Source ↗official
HELM Safetybbq#74 / 800.714Source ↗official
HELM Safetyharmbench#31 / 800.841Source ↗official
HELM Safetysimple_safety_tests#48 / 800.98Source ↗official
HELM Safetyxstest#46 / 800.954Source ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#21 / 270.615Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AILuminate General Purpose AI Chatchild_sexual_exploitation#27 / 329.1Source ↗official
AILuminate General Purpose AI Chatdefamation#23 / 3210.6Source ↗official
AILuminate General Purpose AI Chathate#25 / 324Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#22 / 3218.8Source ↗official
AILuminate General Purpose AI Chatintellectual_property#21 / 328.2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#24 / 329Source ↗official
AILuminate General Purpose AI Chatprivacy#23 / 328.1Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#22 / 329.2Source ↗official
AILuminate General Purpose AI Chatsexual_content#18 / 3221Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#21 / 3213.9Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#22 / 329.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#22 / 3212.7Source ↗official
HELM Safetyanthropic_red_team#50 / 800.993Source ↗official
HELM Safetyharmbench#49 / 800.841Source ↗official
HELM Safetysimple_safety_tests#31 / 800.98Source ↗official
HELM Safetyxstest#46 / 800.954Source ↗official