Model profile
Metadeveloper
2026-07-09release date
#10 / 305overall rank
14eval lineages
Evidence summary
Muse Spark 1.1 has an estimated overall rank of #10; its 90% source-sensitivity interval is #5–#200. Its behavior-only rank is #4; company governance moves the combined estimate to #10. Published evidence spans 14 evals and 7 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (harmful_attack_non_success_rate, #1 of 241); its weakest is SpeciEval (land_animal_4ns, #103 of 105).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗muse-spark-1-1
- OpenRouter ↗meta/muse-spark-1.1
- Evaluation report and model page ↗Exact model document · Meta · first party
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Source |
|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | #58 / 327 | ↓0.4998 | Source ↗official | |
| Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating | #81 / 112 | ↑1420.0 | Source ↗official | |
| CAIS Risk Indexagent_red_teaming | #7 / 44 | ↓44.8 | Source ↗official | |
| CAIS Risk Indexbioweapons_assistance | #4 / 49 | ↓22.3 | Source ↗official | |
| CAIS Risk Indexhle_overconfidence | #2 / 50 | ↓25.3 | Source ↗official | |
| CAIS Risk Indexmachiavelli | #23 / 46 | ↓88.2 | Source ↗official | |
| CAIS Risk Indexmask | #1 / 52 | ↓3.8 | Source ↗official | |
| CAIS Risk Indexpolitical_manipulation | #1 / 33 | ↓27.9 | Source ↗official | |
| CAIS Risk Indextextquests_harm | #10 / 49 | ↓15.2 | Source ↗official | |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | #18 / 241 | ↑49.87 | Source ↗official | |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | #31 / 241 | ↑93.67 | Source ↗official | |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | #1 / 241 | ↑100 | Source ↗official | |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | #135 / 239 | ↑95.64 | Source ↗official | |
| FORTRESSaverage_risk_score | #3 / 49 | ↓12.4 | Source ↗official | |
| FORTRESSover_refusal_score | #33 / 46 | ↓6.92 | Source ↗official | |
| MACHIAVELLIdeception_relative_random_pct | #23 / 50 | ↓88.2 | Source ↗official | |
| Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns | #12 / 19 | ↓12 | Source ↗official | |
| SM-Benchadversarial | #53 / 78 | ↑79.51 | Source ↗official | |
| SM-Benchambiguous_interpretation | #4 / 78 | ↑93.15 | Source ↗official | |
| SM-Benchanti_hallucination | #14 / 78 | ↑98.43 | Source ↗official | |
| SM-Bencheq_boundaries | #26 / 78 | ↑68.26 | Source ↗official | |
| SM-Benchoverfit | #14 / 78 | ↑87.98 | Source ↗official | |
| SpeciEvalbelief_animal_sentience | #1 / 105 | ↑7 | Source ↗official | |
| SpeciEvalland_animal_4ns | #103 / 105 | ↓5.28 | Source ↗official | |
| SpeciEvalsea_animal_4ns | #55 / 105 | ↓4.78 | Source ↗official | |
| SpeciEvalspeciesism | #60 / 105 | ↓2.1 | Source ↗official |