← Models

Model profile

Olmo 2 0325 32B Instruct

Ai2developer
2025-03-13release date
#188 / 267overall rank
7eval lineages

Evidence summary

Olmo 2 0325 32B Instruct has an estimated overall rank of #188; its 90% source-sensitivity interval is #49–#234. Its behavior-only rank is #191; company governance moves the combined estimate to #188. Published evidence spans 7 evals and 6 of 7 behavior components. Its strongest relative result is AILuminate General Purpose AI Chat (child_sexual_exploitation, #6 of 32); its weakest is HELM Safety (bbq, #74 of 80).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AILuminate General Purpose AI Chatchild_sexual_exploitation#6 / 329.1↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#10 / 3210.6↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#8 / 324↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#11 / 3218.8↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#11 / 328.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#9 / 329↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#10 / 328.1↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#11 / 329.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#15 / 3221↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#12 / 3213.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#10 / 329.7↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#11 / 3212.7↓ lowerSource ↗official
HELM Safetyanthropic_red_team#27 / 800.993↑ higherSource ↗official
HELM Safetybbq#74 / 800.714↑ higherSource ↗official
HELM Safetyharmbench#31 / 800.841↑ higherSource ↗official
HELM Safetysimple_safety_tests#48 / 800.98↑ higherSource ↗official
HELM Safetyxstest#46 / 800.954↑ higherSource ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#21 / 270.615↑ higherSource ↗official