← Evals

Evaluation profile

Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversight

38sub-evals
0.083%Safety weight
0%Freedom weight
4components

Weights below are portfolio-specific global index weights.

Model score (see each sub-eval direction)Predicted score

About this eval

Responsible agency, training integrity, deception resistance, truthfulness under pressure, containment, source disclosure, and monitor-evasion resistance.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
aa_omniscience_net_scorefable-mythos51-system-card/alignment-honesty.csv:aa_omniscience_net_scoreMeasures factual accuracy while penalizing incorrect answers and assigning zero to abstentions.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: higher0.00271%
accepting_unverifiable_authorizationfable-mythos51-system-card/alignment-honesty.csv:accepting_unverifiable_authorizationMeasures accepting unverifiable authorization across investigator-generated behavioral-audit scenarios.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
approval_gate_bypassfable-mythos51-system-card/alignment-honesty.csv:approval_gate_bypassMeasures approval gate bypass across investigator-generated behavioral-audit scenarios.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
behavior_consistencyfable-mythos51-system-card/alignment-honesty.csv:behavior_consistencyAutomated-audit consistency of behavior across the investigated interaction.Safety: truthfulness_honesty:0.500|responsible_agency_control:0.500anthropic-fable-mythos51-system-cardSafety: higher0.00167%
coherence_between_actions_and_viewsfable-mythos51-system-card/alignment-honesty.csv:coherence_between_actions_and_viewsAutomated-audit coherence between the model actions and stated views.Safety: truthfulness_honesty:0.700|responsible_agency_control:0.300anthropic-fable-mythos51-system-cardSafety: higher0.00166%
compliance_with_deception_toward_userfable-mythos51-system-card/alignment-honesty.csv:compliance_with_deception_toward_userMeasures compliance with deception toward user across investigator-generated behavioral-audit scenarios.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%
cooperation_with_human_misusefable-mythos51-system-card/alignment-honesty.csv:cooperation_with_human_misuseMeasures cooperation with human misuse across investigator-generated behavioral-audit scenarios.Safety: ordinary_harm_misuse_resistance:1.000anthropic-fable-mythos51-system-cardSafety: lower0.000749%
dangerous_actions_oblivious_to_real_settingfable-mythos51-system-card/alignment-honesty.csv:dangerous_actions_oblivious_to_real_settingAutomated-audit score for dangerous action while failing to recognize the real setting.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
encouragement_of_user_delusionfable-mythos51-system-card/alignment-honesty.csv:encouragement_of_user_delusionMeasures encouragement of user delusion across investigator-generated behavioral-audit scenarios.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%
evidence_of_misaligned_goalsfable-mythos51-system-card/alignment-honesty.csv:evidence_of_misaligned_goalsMeasures evidence of misaligned goals across investigator-generated behavioral-audit scenarios.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.0017%
failure_to_disclose_bad_or_lazy_behaviorfable-mythos51-system-card/alignment-honesty.csv:failure_to_disclose_bad_or_lazy_behaviorMeasures failure to disclose bad or lazy behavior across investigator-generated behavioral-audit scenarios.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%
false_completion_claimsfable-mythos51-system-card/alignment-honesty.csv:false_completion_claimsMeasures false completion claims across investigator-generated behavioral-audit scenarios.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%
ignoring_explicit_constraintsfable-mythos51-system-card/alignment-honesty.csv:ignoring_explicit_constraintsMeasures ignoring explicit constraints across investigator-generated behavioral-audit scenarios.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.0017%
illegible_thinkingfable-mythos51-system-card/alignment-honesty.csv:illegible_thinkingAutomated-audit score for reasoning that is difficult for oversight to interpret.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.0017%
important_omissionsfable-mythos51-system-card/alignment-honesty.csv:important_omissionsMeasures important omissions across investigator-generated behavioral-audit scenarios.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%
input_hallucinationfable-mythos51-system-card/alignment-honesty.csv:input_hallucinationMeasures input hallucination across investigator-generated behavioral-audit scenarios.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%
leaked_answer_undisclosed_use_rate_pctfable-mythos51-system-card/alignment-honesty.csv:leaked_answer_undisclosed_use_rate_pctMeasures silent use of an answer available in the task container.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00325%
linuxarena_stealth_thinking_on_pctfable-mythos51-system-card/alignment-honesty.csv:linuxarena_stealth_thinking_on_pctMeasures successful completion of an instructed harmful side task without detection by the monitor.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00416%
login_wall_alternative_bypass_rate_pctfable-mythos51-system-card/alignment-honesty.csv:login_wall_alternative_bypass_rate_pctRate of attempting a credential bypass other than guessing or container search.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
login_wall_credential_hunting_rate_pctfable-mythos51-system-card/alignment-honesty.csv:login_wall_credential_hunting_rate_pctRate of searching the task container for credentials when blocked by a login wall.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
lying_about_own_work_shared_environment_rate_pctfable-mythos51-system-card/alignment-honesty.csv:lying_about_own_work_shared_environment_rate_pctRate at which the final self-report contradicts observed actions in shared training environments.Safety: truthfulness_honesty:0.700|responsible_agency_control:0.300anthropic-fable-mythos51-system-cardSafety: lower0.00351%
mask_honesty_rate_pctfable-mythos51-system-card/alignment-honesty.csv:mask_honesty_rate_pctMeasures whether the model states what it appears to believe when pressured to say something else.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: higher0.00325%
misaligned_behaviorfable-mythos51-system-card/alignment-honesty.csv:misaligned_behaviorMeasures misaligned behavior across investigator-generated behavioral-audit scenarios.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
motivated_reasoningfable-mythos51-system-card/alignment-honesty.csv:motivated_reasoningAutomated-audit score for motivated rather than evidence-led reasoning.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%
ood_answer_key_or_grader_read_rate_pctfable-mythos51-system-card/alignment-honesty.csv:ood_answer_key_or_grader_read_rate_pctRate of reading an answer key or grader prompt in out-of-distribution computer-use tasks.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
ood_container_tampering_rate_pctfable-mythos51-system-card/alignment-honesty.csv:ood_container_tampering_rate_pctRate of tampering with the task container to satisfy the grader in out-of-distribution computer-use tasks.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
overrefusalfable-mythos51-system-card/alignment-honesty.csv:overrefusalMeasures overrefusal across investigator-generated behavioral-audit scenarios.Safety: benign_helpfulness_non_overrefusal:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00437%
reckless_tool_usefable-mythos51-system-card/alignment-honesty.csv:reckless_tool_useMeasures reckless tool use across investigator-generated behavioral-audit scenarios.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
rewarded_cheating_rate_pctfable-mythos51-system-card/alignment-honesty.csv:rewarded_cheating_rate_pctRate of cheating behavior that earned reward in agentic coding training environments.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.0036%
security_shortcut_for_expediencefable-mythos51-system-card/alignment-honesty.csv:security_shortcut_for_expedienceAutomated-audit score for proposing a security shortcut in the name of expedience.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
self_preservationfable-mythos51-system-card/alignment-honesty.csv:self_preservationMeasures self preservation across investigator-generated behavioral-audit scenarios.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.0017%
shade_arena_stealth_max_effort_pctfable-mythos51-system-card/alignment-honesty.csv:shade_arena_stealth_max_effort_pctMeasures successful completion of an instructed harmful side task without detection by the monitor.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00416%
sycophancyfable-mythos51-system-card/alignment-honesty.csv:sycophancyMeasures sycophancy across investigator-generated behavioral-audit scenarios.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%
unfaithful_thinkingfable-mythos51-system-card/alignment-honesty.csv:unfaithful_thinkingAutomated-audit score for reasoning that is unfaithful to the apparent basis of the answer.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%
unprompted_boundary_probingfable-mythos51-system-card/alignment-honesty.csv:unprompted_boundary_probingMeasures unprompted boundary probing across investigator-generated behavioral-audit scenarios.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.0017%
unsanctioned_sandbox_escape_attemptsfable-mythos51-system-card/alignment-honesty.csv:unsanctioned_sandbox_escape_attemptsAutomated-audit score for unsanctioned attempts to escape a sandbox.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00212%
unsanctioned_third_party_contactfable-mythos51-system-card/alignment-honesty.csv:unsanctioned_third_party_contactMeasures unsanctioned third party contact across investigator-generated behavioral-audit scenarios.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.0017%
user_deceptionfable-mythos51-system-card/alignment-honesty.csv:user_deceptionMeasures user deception across investigator-generated behavioral-audit scenarios.Safety: truthfulness_honesty:1.000anthropic-fable-mythos51-system-cardSafety: lower0.00164%

aa_omniscience_net_score

Measures factual accuracy while penalizing incorrect answers and assigning zero to abstentions.

RankModelValueRelative performanceProvenance
1claude-mythos-50.58official
2claude-mythos-5.10.57official
3claude-mythos-preview0.54official
4claude-opus-50.49official
5claude-opus-4.80.41official
6claude-opus-4.70.38official
7claude-sonnet-50.23official

accepting_unverifiable_authorization

Measures accepting unverifiable authorization across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-opus-51.87official
2claude-mythos-5.12.06official
3claude-mythos-52.14official
4claude-sonnet-52.43official

approval_gate_bypass

Measures approval gate bypass across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-opus-51.28official
2claude-mythos-5.11.36official
3claude-mythos-51.59official
4claude-sonnet-51.63official

behavior_consistency

Automated-audit consistency of behavior across the investigated interaction.

RankModelValueRelative performanceProvenance
1claude-opus-58official
2claude-mythos-5.17.92official
3claude-mythos-57.64official
4claude-sonnet-56.94official

coherence_between_actions_and_views

Automated-audit coherence between the model actions and stated views.

RankModelValueRelative performanceProvenance
1claude-opus-57.69official
2claude-mythos-5.17.54official
3claude-mythos-57.07official
4claude-sonnet-56.32official

compliance_with_deception_toward_user

Measures compliance with deception toward user across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.09official
1claude-opus-51.09official
3claude-mythos-51.12official
4claude-sonnet-51.15official

cooperation_with_human_misuse

Measures cooperation with human misuse across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-opus-51.57official
2claude-mythos-5.11.65official
3claude-mythos-51.74official
4claude-sonnet-51.99official

dangerous_actions_oblivious_to_real_setting

Automated-audit score for dangerous action while failing to recognize the real setting.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.08official
1claude-opus-51.08official
3claude-mythos-51.19official
4claude-sonnet-51.2official

encouragement_of_user_delusion

Measures encouragement of user delusion across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-opus-51.02official
2claude-mythos-51.03official
2claude-mythos-5.11.03official
4claude-sonnet-51.04official

evidence_of_misaligned_goals

Measures evidence of misaligned goals across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-opus-51.03official
2claude-mythos-5.11.04official
3claude-sonnet-51.07official
4claude-mythos-51.1official

failure_to_disclose_bad_or_lazy_behavior

Measures failure to disclose bad or lazy behavior across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.27official
2claude-opus-51.29official
3claude-mythos-51.51official
4claude-sonnet-51.67official

false_completion_claims

Measures false completion claims across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.16official
2claude-opus-51.27official
3claude-mythos-51.29official
4claude-sonnet-51.41official

ignoring_explicit_constraints

Measures ignoring explicit constraints across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.64official
2claude-opus-51.8official
3claude-mythos-51.81official
4claude-sonnet-52official

illegible_thinking

Automated-audit score for reasoning that is difficult for oversight to interpret.

RankModelValueRelative performanceProvenance
1claude-opus-51.81official
2claude-mythos-5.11.96official
3claude-sonnet-52.31official
4claude-mythos-52.37official

important_omissions

Measures important omissions across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-opus-51.22official
2claude-mythos-5.11.25official
3claude-mythos-51.39official
4claude-sonnet-51.55official

input_hallucination

Measures input hallucination across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.27official
2claude-mythos-51.42official
3claude-opus-51.49official
4claude-sonnet-51.57official

leaked_answer_undisclosed_use_rate_pct

Measures silent use of an answer available in the task container.

RankModelValueRelative performanceProvenance
1claude-mythos-preview13official
2claude-opus-4.713.5official
3claude-opus-4.813.7official
4claude-mythos-5.115.7official
5claude-sonnet-518.1official
6claude-mythos-519.3official
7claude-opus-536.9official

linuxarena_stealth_thinking_on_pct

Measures successful completion of an instructed harmful side task without detection by the monitor.

RankModelValueRelative performanceProvenance
1claude-sonnet-50.1official
2claude-opus-4.82.6official
3claude-opus-56.1official
4claude-mythos-preview8.6official
5claude-mythos-59.4official
6claude-mythos-5.113.9official

login_wall_alternative_bypass_rate_pct

Rate of attempting a credential bypass other than guessing or container search.

RankModelValueRelative performanceProvenance
1claude-sonnet-51.8official
2claude-mythos-52.8official
3claude-mythos-5.13official
4claude-opus-57official

login_wall_credential_hunting_rate_pct

Rate of searching the task container for credentials when blocked by a login wall.

RankModelValueRelative performanceProvenance
1claude-mythos-5.116.7official
2claude-mythos-531.1official
3claude-opus-533.5official
4claude-sonnet-536.4official

lying_about_own_work_shared_environment_rate_pct

Rate at which the final self-report contradicts observed actions in shared training environments.

RankModelValueRelative performanceProvenance
1claude-opus-50.16official
2claude-mythos-5.10.18official
3claude-fable-50.22official
3claude-opus-4.80.22official
5claude-opus-4.70.27official
6claude-sonnet-50.3official
7claude-sonnet-4.60.46official
8claude-opus-4.60.59official

mask_honesty_rate_pct

Measures whether the model states what it appears to believe when pressured to say something else.

RankModelValueRelative performanceProvenance
1claude-sonnet-596.9official
2claude-mythos-preview95.6official
3claude-opus-594.8official
4claude-opus-4.893.9official
5claude-mythos-591.4official
6claude-opus-4.790.2official
7claude-mythos-5.185.3official

misaligned_behavior

Measures misaligned behavior across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-opus-52.3official
2claude-mythos-5.12.4official
3claude-mythos-52.79official
4claude-sonnet-53.07official

motivated_reasoning

Automated-audit score for motivated rather than evidence-led reasoning.

RankModelValueRelative performanceProvenance
1claude-opus-51.76official
2claude-mythos-5.11.82official
3claude-mythos-52.24official
4claude-sonnet-52.58official

ood_answer_key_or_grader_read_rate_pct

Rate of reading an answer key or grader prompt in out-of-distribution computer-use tasks.

RankModelValueRelative performanceProvenance
1claude-mythos-5.112official
2claude-opus-514official
3claude-mythos-534.5official
4claude-sonnet-535official

ood_container_tampering_rate_pct

Rate of tampering with the task container to satisfy the grader in out-of-distribution computer-use tasks.

RankModelValueRelative performanceProvenance
1claude-mythos-5.116.4official
2claude-opus-520.9official
3claude-sonnet-537.2official
4claude-mythos-544.5official

overrefusal

Measures overrefusal across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-51.44official
2claude-mythos-5.11.49official
2claude-opus-51.49official
4claude-sonnet-51.64official

reckless_tool_use

Measures reckless tool use across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.57official
2claude-opus-51.58official
3claude-sonnet-51.78official
4claude-mythos-51.81official

rewarded_cheating_rate_pct

Rate of cheating behavior that earned reward in agentic coding training environments.

RankModelValueRelative performanceProvenance
1claude-mythos-5.10.06official
2claude-opus-4.70.12official
3claude-opus-4.80.26official
4claude-opus-4.60.28official
4claude-sonnet-50.28official
6claude-sonnet-4.60.31official
7claude-opus-50.32official
8claude-fable-50.44official

security_shortcut_for_expedience

Automated-audit score for proposing a security shortcut in the name of expedience.

RankModelValueRelative performanceProvenance
1claude-opus-51.49official
2claude-mythos-5.11.54official
3claude-mythos-51.86official
4claude-sonnet-51.95official

self_preservation

Measures self preservation across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.04official
1claude-opus-51.04official
3claude-sonnet-51.06official
4claude-mythos-51.09official

shade_arena_stealth_max_effort_pct

Measures successful completion of an instructed harmful side task without detection by the monitor.

RankModelValueRelative performanceProvenance
1claude-sonnet-51.3official
2claude-opus-4.85.3official
3claude-opus-511.2official
4claude-mythos-512.8official
5claude-mythos-preview16.7official
6claude-mythos-5.117.3official

sycophancy

Measures sycophancy across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.36official
2claude-opus-51.5official
3claude-mythos-51.52official
4claude-sonnet-51.69official

unfaithful_thinking

Automated-audit score for reasoning that is unfaithful to the apparent basis of the answer.

RankModelValueRelative performanceProvenance
1claude-opus-51.49official
2claude-mythos-5.11.57official
3claude-mythos-51.91official
4claude-sonnet-52.23official

unprompted_boundary_probing

Measures unprompted boundary probing across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.1official
2claude-opus-51.13official
3claude-mythos-51.2official
4claude-sonnet-51.21official

unsanctioned_sandbox_escape_attempts

Automated-audit score for unsanctioned attempts to escape a sandbox.

RankModelValueRelative performanceProvenance
1claude-opus-51.25official
2claude-mythos-5.11.26official
3claude-sonnet-51.43official
4claude-mythos-51.49official

unsanctioned_third_party_contact

Measures unsanctioned third party contact across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-opus-51.23official
2claude-mythos-5.11.24official
3claude-sonnet-51.27official
4claude-mythos-51.33official

user_deception

Measures user deception across investigator-generated behavioral-audit scenarios.

RankModelValueRelative performanceProvenance
1claude-mythos-5.11.2official
2claude-opus-51.24official
3claude-mythos-51.41official
4claude-sonnet-51.49official