Developer
Microsoft
22 indexed models; 7 currently meet the evidence threshold for the overall ranking. Together they have results from 17 evaluations.
Models by Microsoft
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Phi 3.5 Moe Instruct | 5 | 6/7 | 33 | 2024-08-22 |
| Phi 4 Reasoning Plus | 3 | 5/7 | 125 | 2025-04-17 |
| Phi 3 Mini 4K Instruct | 4 | 3/7 | 137 | 2024-04-22 |
| Phi 4 | 6 | 5/7 | 140 | 2024-12-11 |
| Phi 3.5 Mini Instruct | 5 | 6/7 | 225 | 2024-08-22 |
| Phi 2 | 3 | 3/7 | 238 | — |
| Phi 4 Mini | 4 | 4/7 | 280 | — |
| Orca 2 13B | 1 | 1/7 | — | 2023-11-14 |
| Orca 2 7B | 1 | 1/7 | — | 2023-11-14 |
| Phi 1 5 | 1 | 1/7 | — | — |
| Phi 3 Medium | 1 | 2/7 | — | — |
| Phi 3 Medium 128K Instruct | 2 | 3/7 | — | — |
| Phi 3 Medium 4K Instruct | 1 | 3/7 | — | — |
| Phi 3 Mini | 1 | 2/7 | — | — |
| Phi 3 Mini 128K Instruct | 1 | 3/7 | — | — |
| Phi 3 Small | 1 | 2/7 | — | — |
| Phi 3 Small 128K Instruct | 1 | 3/7 | — | — |
| Phi 3 Small 8K Instruct | 1 | 3/7 | — | — |
| Phi-4 Multimodal Instruct | 1 | 1/7 | — | 2025-02-26 |
| Wizardlm 13B | 1 | 3/7 | — | 2023-05-13 |
| Wizardlm 2 8X22B | 2 | 3/7 | — | — |
| Wizardlm 7B | 1 | 3/7 | — | 2023-04-23 |
Evaluations covering Microsoft models (17)
AA-Omniscience · AgentDrive Safety Compliance · AILuminate General Purpose AI Chat · Cisco AI Defense Rolling Single-Turn Leaderboard · Confabulations · DSPSafeBench · Enkrypt AI Safety Leaderboard · FlagEval Safety and Values · HalluVerse-M3 Hallucination Recognition · HarmBench · Large-scale Moral Machine experiment on LLMs · Microsoft Phi Safety Panels · Open LLM Safety Index · PandaBench JBB direct-request panel · SafetyBench · UAVBench safety-critical decision recognition · Vectara HHEM Factual Consistency