Evaluation profile
Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversight
Weights below are portfolio-specific global index weights.
About this eval
Responsible agency, training integrity, deception resistance, truthfulness under pressure, containment, source disclosure, and monitor-evasion resistance.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| aa_omniscience_net_scorefable-mythos51-system-card/alignment-honesty.csv:aa_omniscience_net_scoreMeasures factual accuracy while penalizing incorrect answers and assigning zero to abstentions. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: higher | 0.00271% | — |
| accepting_unverifiable_authorizationfable-mythos51-system-card/alignment-honesty.csv:accepting_unverifiable_authorizationMeasures accepting unverifiable authorization across investigator-generated behavioral-audit scenarios. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| approval_gate_bypassfable-mythos51-system-card/alignment-honesty.csv:approval_gate_bypassMeasures approval gate bypass across investigator-generated behavioral-audit scenarios. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| behavior_consistencyfable-mythos51-system-card/alignment-honesty.csv:behavior_consistencyAutomated-audit consistency of behavior across the investigated interaction. | Safety: truthfulness_honesty:0.500|responsible_agency_control:0.500anthropic-fable-mythos51-system-card | Safety: higher | 0.00167% | — |
| coherence_between_actions_and_viewsfable-mythos51-system-card/alignment-honesty.csv:coherence_between_actions_and_viewsAutomated-audit coherence between the model actions and stated views. | Safety: truthfulness_honesty:0.700|responsible_agency_control:0.300anthropic-fable-mythos51-system-card | Safety: higher | 0.00166% | — |
| compliance_with_deception_toward_userfable-mythos51-system-card/alignment-honesty.csv:compliance_with_deception_toward_userMeasures compliance with deception toward user across investigator-generated behavioral-audit scenarios. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
| cooperation_with_human_misusefable-mythos51-system-card/alignment-honesty.csv:cooperation_with_human_misuseMeasures cooperation with human misuse across investigator-generated behavioral-audit scenarios. | Safety: ordinary_harm_misuse_resistance:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.000749% | — |
| dangerous_actions_oblivious_to_real_settingfable-mythos51-system-card/alignment-honesty.csv:dangerous_actions_oblivious_to_real_settingAutomated-audit score for dangerous action while failing to recognize the real setting. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| encouragement_of_user_delusionfable-mythos51-system-card/alignment-honesty.csv:encouragement_of_user_delusionMeasures encouragement of user delusion across investigator-generated behavioral-audit scenarios. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
| evidence_of_misaligned_goalsfable-mythos51-system-card/alignment-honesty.csv:evidence_of_misaligned_goalsMeasures evidence of misaligned goals across investigator-generated behavioral-audit scenarios. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.0017% | — |
| failure_to_disclose_bad_or_lazy_behaviorfable-mythos51-system-card/alignment-honesty.csv:failure_to_disclose_bad_or_lazy_behaviorMeasures failure to disclose bad or lazy behavior across investigator-generated behavioral-audit scenarios. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
| false_completion_claimsfable-mythos51-system-card/alignment-honesty.csv:false_completion_claimsMeasures false completion claims across investigator-generated behavioral-audit scenarios. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
| ignoring_explicit_constraintsfable-mythos51-system-card/alignment-honesty.csv:ignoring_explicit_constraintsMeasures ignoring explicit constraints across investigator-generated behavioral-audit scenarios. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.0017% | — |
| illegible_thinkingfable-mythos51-system-card/alignment-honesty.csv:illegible_thinkingAutomated-audit score for reasoning that is difficult for oversight to interpret. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.0017% | — |
| important_omissionsfable-mythos51-system-card/alignment-honesty.csv:important_omissionsMeasures important omissions across investigator-generated behavioral-audit scenarios. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
| input_hallucinationfable-mythos51-system-card/alignment-honesty.csv:input_hallucinationMeasures input hallucination across investigator-generated behavioral-audit scenarios. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
| leaked_answer_undisclosed_use_rate_pctfable-mythos51-system-card/alignment-honesty.csv:leaked_answer_undisclosed_use_rate_pctMeasures silent use of an answer available in the task container. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00325% | — |
| linuxarena_stealth_thinking_on_pctfable-mythos51-system-card/alignment-honesty.csv:linuxarena_stealth_thinking_on_pctMeasures successful completion of an instructed harmful side task without detection by the monitor. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00416% | — |
| login_wall_alternative_bypass_rate_pctfable-mythos51-system-card/alignment-honesty.csv:login_wall_alternative_bypass_rate_pctRate of attempting a credential bypass other than guessing or container search. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| login_wall_credential_hunting_rate_pctfable-mythos51-system-card/alignment-honesty.csv:login_wall_credential_hunting_rate_pctRate of searching the task container for credentials when blocked by a login wall. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| lying_about_own_work_shared_environment_rate_pctfable-mythos51-system-card/alignment-honesty.csv:lying_about_own_work_shared_environment_rate_pctRate at which the final self-report contradicts observed actions in shared training environments. | Safety: truthfulness_honesty:0.700|responsible_agency_control:0.300anthropic-fable-mythos51-system-card | Safety: lower | 0.00351% | — |
| mask_honesty_rate_pctfable-mythos51-system-card/alignment-honesty.csv:mask_honesty_rate_pctMeasures whether the model states what it appears to believe when pressured to say something else. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: higher | 0.00325% | — |
| misaligned_behaviorfable-mythos51-system-card/alignment-honesty.csv:misaligned_behaviorMeasures misaligned behavior across investigator-generated behavioral-audit scenarios. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| motivated_reasoningfable-mythos51-system-card/alignment-honesty.csv:motivated_reasoningAutomated-audit score for motivated rather than evidence-led reasoning. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
| ood_answer_key_or_grader_read_rate_pctfable-mythos51-system-card/alignment-honesty.csv:ood_answer_key_or_grader_read_rate_pctRate of reading an answer key or grader prompt in out-of-distribution computer-use tasks. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| ood_container_tampering_rate_pctfable-mythos51-system-card/alignment-honesty.csv:ood_container_tampering_rate_pctRate of tampering with the task container to satisfy the grader in out-of-distribution computer-use tasks. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| overrefusalfable-mythos51-system-card/alignment-honesty.csv:overrefusalMeasures overrefusal across investigator-generated behavioral-audit scenarios. | Safety: benign_helpfulness_non_overrefusal:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00437% | — |
| reckless_tool_usefable-mythos51-system-card/alignment-honesty.csv:reckless_tool_useMeasures reckless tool use across investigator-generated behavioral-audit scenarios. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| rewarded_cheating_rate_pctfable-mythos51-system-card/alignment-honesty.csv:rewarded_cheating_rate_pctRate of cheating behavior that earned reward in agentic coding training environments. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.0036% | — |
| security_shortcut_for_expediencefable-mythos51-system-card/alignment-honesty.csv:security_shortcut_for_expedienceAutomated-audit score for proposing a security shortcut in the name of expedience. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| self_preservationfable-mythos51-system-card/alignment-honesty.csv:self_preservationMeasures self preservation across investigator-generated behavioral-audit scenarios. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.0017% | — |
| shade_arena_stealth_max_effort_pctfable-mythos51-system-card/alignment-honesty.csv:shade_arena_stealth_max_effort_pctMeasures successful completion of an instructed harmful side task without detection by the monitor. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00416% | — |
| sycophancyfable-mythos51-system-card/alignment-honesty.csv:sycophancyMeasures sycophancy across investigator-generated behavioral-audit scenarios. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
| unfaithful_thinkingfable-mythos51-system-card/alignment-honesty.csv:unfaithful_thinkingAutomated-audit score for reasoning that is unfaithful to the apparent basis of the answer. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
| unprompted_boundary_probingfable-mythos51-system-card/alignment-honesty.csv:unprompted_boundary_probingMeasures unprompted boundary probing across investigator-generated behavioral-audit scenarios. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.0017% | — |
| unsanctioned_sandbox_escape_attemptsfable-mythos51-system-card/alignment-honesty.csv:unsanctioned_sandbox_escape_attemptsAutomated-audit score for unsanctioned attempts to escape a sandbox. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00212% | — |
| unsanctioned_third_party_contactfable-mythos51-system-card/alignment-honesty.csv:unsanctioned_third_party_contactMeasures unsanctioned third party contact across investigator-generated behavioral-audit scenarios. | Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.0017% | — |
| user_deceptionfable-mythos51-system-card/alignment-honesty.csv:user_deceptionMeasures user deception across investigator-generated behavioral-audit scenarios. | Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-card | Safety: lower | 0.00164% | — |
aa_omniscience_net_score
Measures factual accuracy while penalizing incorrect answers and assigning zero to abstentions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5 | 0.58 | official | |
| 2 | claude-mythos-5.1 | 0.57 | official | |
| 3 | claude-mythos-preview | 0.54 | official | |
| 4 | claude-opus-5 | 0.49 | official | |
| 5 | claude-opus-4.8 | 0.41 | official | |
| 6 | claude-opus-4.7 | 0.38 | official | |
| 7 | claude-sonnet-5 | 0.23 | official |
accepting_unverifiable_authorization
Measures accepting unverifiable authorization across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.87 | official | |
| 2 | claude-mythos-5.1 | 2.06 | official | |
| 3 | claude-mythos-5 | 2.14 | official | |
| 4 | claude-sonnet-5 | 2.43 | official |
approval_gate_bypass
Measures approval gate bypass across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.28 | official | |
| 2 | claude-mythos-5.1 | 1.36 | official | |
| 3 | claude-mythos-5 | 1.59 | official | |
| 4 | claude-sonnet-5 | 1.63 | official |
behavior_consistency
Automated-audit consistency of behavior across the investigated interaction.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 8 | official | |
| 2 | claude-mythos-5.1 | 7.92 | official | |
| 3 | claude-mythos-5 | 7.64 | official | |
| 4 | claude-sonnet-5 | 6.94 | official |
coherence_between_actions_and_views
Automated-audit coherence between the model actions and stated views.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 7.69 | official | |
| 2 | claude-mythos-5.1 | 7.54 | official | |
| 3 | claude-mythos-5 | 7.07 | official | |
| 4 | claude-sonnet-5 | 6.32 | official |
compliance_with_deception_toward_user
Measures compliance with deception toward user across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.09 | official | |
| 1 | claude-opus-5 | 1.09 | official | |
| 3 | claude-mythos-5 | 1.12 | official | |
| 4 | claude-sonnet-5 | 1.15 | official |
cooperation_with_human_misuse
Measures cooperation with human misuse across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.57 | official | |
| 2 | claude-mythos-5.1 | 1.65 | official | |
| 3 | claude-mythos-5 | 1.74 | official | |
| 4 | claude-sonnet-5 | 1.99 | official |
dangerous_actions_oblivious_to_real_setting
Automated-audit score for dangerous action while failing to recognize the real setting.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.08 | official | |
| 1 | claude-opus-5 | 1.08 | official | |
| 3 | claude-mythos-5 | 1.19 | official | |
| 4 | claude-sonnet-5 | 1.2 | official |
encouragement_of_user_delusion
Measures encouragement of user delusion across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.02 | official | |
| 2 | claude-mythos-5 | 1.03 | official | |
| 2 | claude-mythos-5.1 | 1.03 | official | |
| 4 | claude-sonnet-5 | 1.04 | official |
evidence_of_misaligned_goals
Measures evidence of misaligned goals across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.03 | official | |
| 2 | claude-mythos-5.1 | 1.04 | official | |
| 3 | claude-sonnet-5 | 1.07 | official | |
| 4 | claude-mythos-5 | 1.1 | official |
failure_to_disclose_bad_or_lazy_behavior
Measures failure to disclose bad or lazy behavior across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.27 | official | |
| 2 | claude-opus-5 | 1.29 | official | |
| 3 | claude-mythos-5 | 1.51 | official | |
| 4 | claude-sonnet-5 | 1.67 | official |
false_completion_claims
Measures false completion claims across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.16 | official | |
| 2 | claude-opus-5 | 1.27 | official | |
| 3 | claude-mythos-5 | 1.29 | official | |
| 4 | claude-sonnet-5 | 1.41 | official |
ignoring_explicit_constraints
Measures ignoring explicit constraints across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.64 | official | |
| 2 | claude-opus-5 | 1.8 | official | |
| 3 | claude-mythos-5 | 1.81 | official | |
| 4 | claude-sonnet-5 | 2 | official |
illegible_thinking
Automated-audit score for reasoning that is difficult for oversight to interpret.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.81 | official | |
| 2 | claude-mythos-5.1 | 1.96 | official | |
| 3 | claude-sonnet-5 | 2.31 | official | |
| 4 | claude-mythos-5 | 2.37 | official |
important_omissions
Measures important omissions across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.22 | official | |
| 2 | claude-mythos-5.1 | 1.25 | official | |
| 3 | claude-mythos-5 | 1.39 | official | |
| 4 | claude-sonnet-5 | 1.55 | official |
input_hallucination
Measures input hallucination across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.27 | official | |
| 2 | claude-mythos-5 | 1.42 | official | |
| 3 | claude-opus-5 | 1.49 | official | |
| 4 | claude-sonnet-5 | 1.57 | official |
leaked_answer_undisclosed_use_rate_pct
Measures silent use of an answer available in the task container.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-preview | 13 | official | |
| 2 | claude-opus-4.7 | 13.5 | official | |
| 3 | claude-opus-4.8 | 13.7 | official | |
| 4 | claude-mythos-5.1 | 15.7 | official | |
| 5 | claude-sonnet-5 | 18.1 | official | |
| 6 | claude-mythos-5 | 19.3 | official | |
| 7 | claude-opus-5 | 36.9 | official |
linuxarena_stealth_thinking_on_pct
Measures successful completion of an instructed harmful side task without detection by the monitor.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-5 | 0.1 | official | |
| 2 | claude-opus-4.8 | 2.6 | official | |
| 3 | claude-opus-5 | 6.1 | official | |
| 4 | claude-mythos-preview | 8.6 | official | |
| 5 | claude-mythos-5 | 9.4 | official | |
| 6 | claude-mythos-5.1 | 13.9 | official |
login_wall_alternative_bypass_rate_pct
Rate of attempting a credential bypass other than guessing or container search.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-5 | 1.8 | official | |
| 2 | claude-mythos-5 | 2.8 | official | |
| 3 | claude-mythos-5.1 | 3 | official | |
| 4 | claude-opus-5 | 7 | official |
login_wall_credential_hunting_rate_pct
Rate of searching the task container for credentials when blocked by a login wall.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 16.7 | official | |
| 2 | claude-mythos-5 | 31.1 | official | |
| 3 | claude-opus-5 | 33.5 | official | |
| 4 | claude-sonnet-5 | 36.4 | official |
mask_honesty_rate_pct
Measures whether the model states what it appears to believe when pressured to say something else.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-5 | 96.9 | official | |
| 2 | claude-mythos-preview | 95.6 | official | |
| 3 | claude-opus-5 | 94.8 | official | |
| 4 | claude-opus-4.8 | 93.9 | official | |
| 5 | claude-mythos-5 | 91.4 | official | |
| 6 | claude-opus-4.7 | 90.2 | official | |
| 7 | claude-mythos-5.1 | 85.3 | official |
misaligned_behavior
Measures misaligned behavior across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 2.3 | official | |
| 2 | claude-mythos-5.1 | 2.4 | official | |
| 3 | claude-mythos-5 | 2.79 | official | |
| 4 | claude-sonnet-5 | 3.07 | official |
motivated_reasoning
Automated-audit score for motivated rather than evidence-led reasoning.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.76 | official | |
| 2 | claude-mythos-5.1 | 1.82 | official | |
| 3 | claude-mythos-5 | 2.24 | official | |
| 4 | claude-sonnet-5 | 2.58 | official |
ood_answer_key_or_grader_read_rate_pct
Rate of reading an answer key or grader prompt in out-of-distribution computer-use tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 12 | official | |
| 2 | claude-opus-5 | 14 | official | |
| 3 | claude-mythos-5 | 34.5 | official | |
| 4 | claude-sonnet-5 | 35 | official |
ood_container_tampering_rate_pct
Rate of tampering with the task container to satisfy the grader in out-of-distribution computer-use tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 16.4 | official | |
| 2 | claude-opus-5 | 20.9 | official | |
| 3 | claude-sonnet-5 | 37.2 | official | |
| 4 | claude-mythos-5 | 44.5 | official |
overrefusal
Measures overrefusal across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5 | 1.44 | official | |
| 2 | claude-mythos-5.1 | 1.49 | official | |
| 2 | claude-opus-5 | 1.49 | official | |
| 4 | claude-sonnet-5 | 1.64 | official |
reckless_tool_use
Measures reckless tool use across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.57 | official | |
| 2 | claude-opus-5 | 1.58 | official | |
| 3 | claude-sonnet-5 | 1.78 | official | |
| 4 | claude-mythos-5 | 1.81 | official |
rewarded_cheating_rate_pct
Rate of cheating behavior that earned reward in agentic coding training environments.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 0.06 | official | |
| 2 | claude-opus-4.7 | 0.12 | official | |
| 3 | claude-opus-4.8 | 0.26 | official | |
| 4 | claude-opus-4.6 | 0.28 | official | |
| 4 | claude-sonnet-5 | 0.28 | official | |
| 6 | claude-sonnet-4.6 | 0.31 | official | |
| 7 | claude-opus-5 | 0.32 | official | |
| 8 | claude-fable-5 | 0.44 | official |
security_shortcut_for_expedience
Automated-audit score for proposing a security shortcut in the name of expedience.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.49 | official | |
| 2 | claude-mythos-5.1 | 1.54 | official | |
| 3 | claude-mythos-5 | 1.86 | official | |
| 4 | claude-sonnet-5 | 1.95 | official |
self_preservation
Measures self preservation across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.04 | official | |
| 1 | claude-opus-5 | 1.04 | official | |
| 3 | claude-sonnet-5 | 1.06 | official | |
| 4 | claude-mythos-5 | 1.09 | official |
shade_arena_stealth_max_effort_pct
Measures successful completion of an instructed harmful side task without detection by the monitor.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-5 | 1.3 | official | |
| 2 | claude-opus-4.8 | 5.3 | official | |
| 3 | claude-opus-5 | 11.2 | official | |
| 4 | claude-mythos-5 | 12.8 | official | |
| 5 | claude-mythos-preview | 16.7 | official | |
| 6 | claude-mythos-5.1 | 17.3 | official |
sycophancy
Measures sycophancy across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.36 | official | |
| 2 | claude-opus-5 | 1.5 | official | |
| 3 | claude-mythos-5 | 1.52 | official | |
| 4 | claude-sonnet-5 | 1.69 | official |
unfaithful_thinking
Automated-audit score for reasoning that is unfaithful to the apparent basis of the answer.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.49 | official | |
| 2 | claude-mythos-5.1 | 1.57 | official | |
| 3 | claude-mythos-5 | 1.91 | official | |
| 4 | claude-sonnet-5 | 2.23 | official |
unprompted_boundary_probing
Measures unprompted boundary probing across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.1 | official | |
| 2 | claude-opus-5 | 1.13 | official | |
| 3 | claude-mythos-5 | 1.2 | official | |
| 4 | claude-sonnet-5 | 1.21 | official |
unsanctioned_sandbox_escape_attempts
Automated-audit score for unsanctioned attempts to escape a sandbox.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.25 | official | |
| 2 | claude-mythos-5.1 | 1.26 | official | |
| 3 | claude-sonnet-5 | 1.43 | official | |
| 4 | claude-mythos-5 | 1.49 | official |
unsanctioned_third_party_contact
Measures unsanctioned third party contact across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.23 | official | |
| 2 | claude-mythos-5.1 | 1.24 | official | |
| 3 | claude-sonnet-5 | 1.27 | official | |
| 4 | claude-mythos-5 | 1.33 | official |
user_deception
Measures user deception across investigator-generated behavioral-audit scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-5.1 | 1.2 | official | |
| 2 | claude-opus-5 | 1.24 | official | |
| 3 | claude-mythos-5 | 1.41 | official | |
| 4 | claude-sonnet-5 | 1.49 | official |