← Models

Model profile

Muse Spark 1.1

Metadeveloper
2026-07-09release date
#10 / 305overall rank
14eval lineages

Evidence summary

Muse Spark 1.1 has an estimated overall rank of #10; its 90% source-sensitivity interval is #5–#200. Its behavior-only rank is #4; company governance moves the combined estimate to #10. Published evidence spans 14 evals and 7 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (harmful_attack_non_success_rate, #1 of 241); its weakest is SpeciEval (land_animal_4ns, #103 of 105).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#58 / 3270.4998Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#81 / 1121420.0Source ↗official
CAIS Risk Indexagent_red_teaming#7 / 4444.8Source ↗official
CAIS Risk Indexbioweapons_assistance#4 / 4922.3Source ↗official
CAIS Risk Indexhle_overconfidence#2 / 5025.3Source ↗official
CAIS Risk Indexmachiavelli#23 / 4688.2Source ↗official
CAIS Risk Indexmask#1 / 523.8Source ↗official
CAIS Risk Indexpolitical_manipulation#1 / 3327.9Source ↗official
CAIS Risk Indextextquests_harm#10 / 4915.2Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#18 / 24149.87Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#31 / 24193.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#1 / 241100Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#135 / 23995.64Source ↗official
FORTRESSaverage_risk_score#3 / 4912.4Source ↗official
FORTRESSover_refusal_score#33 / 466.92Source ↗official
MACHIAVELLIdeception_relative_random_pct#23 / 5088.2Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#12 / 1912Source ↗official
SM-Benchadversarial#53 / 7879.51Source ↗official
SM-Benchambiguous_interpretation#4 / 7893.15Source ↗official
SM-Benchanti_hallucination#14 / 7898.43Source ↗official
SM-Bencheq_boundaries#26 / 7868.26Source ↗official
SM-Benchoverfit#14 / 7887.98Source ↗official
SpeciEvalbelief_animal_sentience#1 / 1057Source ↗official
SpeciEvalland_animal_4ns#103 / 1055.28Source ↗official
SpeciEvalsea_animal_4ns#55 / 1054.78Source ↗official
SpeciEvalspeciesism#60 / 1052.1Source ↗official