Model profile
Olmo 7B 0724 Instruct
2024-07-09release date
2eval lineages
Evidence summary
Published evidence spans 2 evals and 5 of 7 behavior components. Its strongest relative result is AbstentionBench (stale_f1, #10 of 20); its weakest is AILuminate General Purpose AI Chat (child_sexual_exploitation, #32 of 32).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Hugging Face ↗allenai/OLMo-7B-0724-Instruct-hf
- Official model page ↗Exact model document · Reviewed official Hugging Face owner · official repository
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Better | Source |
|---|---|---|---|---|---|
| AbstentionBenchanswer_unknown_f1 | #16 / 20 | 0.8086 | ↑ higher | Source ↗official | |
| AbstentionBenchfalse_premise_f1 | #13 / 20 | 0.6608 | ↑ higher | Source ↗official | |
| AbstentionBenchstale_f1 | #10 / 20 | 0.6392 | ↑ higher | Source ↗official | |
| AbstentionBenchsubjective_f1 | #16 / 20 | 0.6598 | ↑ higher | Source ↗official | |
| AbstentionBenchunderspecified_context_f1 | #15 / 20 | 0.5001 | ↑ higher | Source ↗official | |
| AbstentionBenchunderspecified_intent_f1 | #16 / 20 | 0.6654 | ↑ higher | Source ↗official | |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | #32 / 32 | 81.2 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatdefamation | #32 / 32 | 74 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chathate | #32 / 32 | 65.6 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | #32 / 32 | 79.2 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatintellectual_property | #32 / 32 | 53.2 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatnon_violent_crimes | #32 / 32 | 75 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatprivacy | #32 / 32 | 62.8 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatsex_related_crimes | #32 / 32 | 75.9 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatsexual_content | #32 / 32 | 70.5 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatspecialized_advice | #32 / 32 | 63.4 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatsuicide_self_harm | #32 / 32 | 70.7 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatviolent_crimes | #32 / 32 | 77.6 | ↓ lower | Source ↗official |