Developer
Ai2
26 indexed models; 5 currently meet the evidence threshold for the overall ranking. Together they have results from 10 evaluations.
Models by Ai2
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Llama 3.1 Tulu 3 8B | 3 | 5/7 | 40 | 2024-11-20 |
| Olmo 2 0325 32B Instruct | 7 | 6/7 | 188 | 2025-03-13 |
| Olmo 2 1124 13B Instruct | 6 | 5/7 | 231 | 2024-11-26 |
| Olmo 2 1124 7B Instruct | 6 | 5/7 | 258 | 2024-11-26 |
| Olmoe 1B 7B 0125 Instruct | 5 | 3/7 | 263 | 2025-01-27 |
| Llama 3.1 Tulu 3 70B | 1 | 1/7 | — | — |
| Llama 3.1 Tulu 3 70B Dpo | 1 | 2/7 | — | 2024-11-20 |
| Llama 3.1 Tulu 3 70B Ppo Rlvf | 1 | 2/7 | — | 2024-11-20 |
| Llama 3.1 Tulu 3 70B Sft | 1 | 2/7 | — | 2024-11-18 |
| Llama 3.1 Tulu 3 8B Dpo | 2 | 5/7 | — | 2024-11-20 |
| Llama 3.1 Tulu 3 8B Ppo Rlvf | 1 | 2/7 | — | 2024-11-20 |
| Llama 3.1 Tulu 3 8B Sft | 2 | 5/7 | — | 2024-11-18 |
| Molmo2 8B | 1 | 1/7 | — | 2025-12-14 |
| Olmo 3 32B Think | 1 | 1/7 | — | 2025-11-19 |
| Olmo 3 7B Instruct | 1 | 1/7 | — | 2025-11-19 |
| Olmo 3 7B Think | 1 | 1/7 | — | 2025-11-18 |
| Olmo 3.1 32B Instruct | 1 | 1/7 | — | 2025-12-10 |
| Olmo 3.1 32B Think | 2 | 2/7 | — | 2025-12-10 |
| Olmo 7B 0724 Instruct | 2 | 5/7 | — | 2024-07-09 |
| Olmo 7B Instruct Hf | 1 | 3/7 | — | — |
| Olmoe 1B 7B 0924 Instruct | 1 | 3/7 | — | — |
| Tulu 2 13B | 1 | 2/7 | — | — |
| Tulu 2 7B | 1 | 2/7 | — | — |
| Tulu 2 Dpo 13B | 1 | 4/7 | — | 2023-11-13 |
| Tulu 2 Dpo 70B | 1 | 4/7 | — | 2023-11-12 |
| Tulu 2 Dpo 7B | 1 | 4/7 | — | 2023-11-13 |
Evaluations covering Ai2 models (10)
AA-Omniscience · AbstentionBench · AILuminate General Purpose AI Chat · DecodingTrust · Enkrypt AI Safety Leaderboard · HELM Safety · HUMAINE Trust, Ethics and Safety · PandaBench JBB direct-request panel · SALAD-Bench · UAVBench safety-critical decision recognition