← Models

Model profile

OLMo 7B Instruct

Ai2developer
2024-07-09release date
Not rankedSafety rank
#117 / 645Freedom rank

Evidence summary

Safety. OLMo 7B Instruct does not meet the evidence gate for a Safety rank. Published Safety evidence spans 2 eval lineages and 5 of 7 components. Its strongest relative result is AbstentionBench (stale_f1, #10 of 20); its weakest is AILuminate General Purpose AI Chat (child_sexual_exploitation, #32 of 32).

Freedom. OLMo 7B Instruct has an estimated Freedom rank of #117; its 90% source-sensitivity interval is #1–#551. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is AILuminate General Purpose AI Chat (child_sexual_exploitation, #1 of 32); its weakest is AILuminate General Purpose AI Chat (child_sexual_exploitation, #1 of 32).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AbstentionBenchanswer_unknown_f1#16 / 200.8086Source ↗official
AbstentionBenchfalse_premise_f1#13 / 200.6608Source ↗official
AbstentionBenchstale_f1#10 / 200.6392Source ↗official
AbstentionBenchsubjective_f1#16 / 200.6598Source ↗official
AbstentionBenchunderspecified_context_f1#15 / 200.5001Source ↗official
AbstentionBenchunderspecified_intent_f1#16 / 200.6654Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#32 / 3281.2Source ↗official
AILuminate General Purpose AI Chatdefamation#32 / 3274Source ↗official
AILuminate General Purpose AI Chathate#32 / 3265.6Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#32 / 3279.2Source ↗official
AILuminate General Purpose AI Chatintellectual_property#32 / 3253.2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#32 / 3275Source ↗official
AILuminate General Purpose AI Chatprivacy#32 / 3262.8Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#32 / 3275.9Source ↗official
AILuminate General Purpose AI Chatsexual_content#32 / 3270.5Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#32 / 3263.4Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#32 / 3270.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#32 / 3277.6Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AILuminate General Purpose AI Chatchild_sexual_exploitation#1 / 3281.2Source ↗official
AILuminate General Purpose AI Chatdefamation#1 / 3274Source ↗official
AILuminate General Purpose AI Chathate#1 / 3265.6Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#1 / 3279.2Source ↗official
AILuminate General Purpose AI Chatintellectual_property#1 / 3253.2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#1 / 3275Source ↗official
AILuminate General Purpose AI Chatprivacy#1 / 3262.8Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#1 / 3275.9Source ↗official
AILuminate General Purpose AI Chatsexual_content#1 / 3270.5Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#1 / 3263.4Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#1 / 3270.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#1 / 3277.6Source ↗official