← Evals

Evaluation profile

UK AISI cyber-evaluation cheating and prompted self-report

2sub-evals
0.168%total index weight
2components

Within-component eval weight: Responsible agency 0.598% · Truthfulness 0.523%.

Model score (see each sub-eval direction)Predicted score

About this eval

Propensity to use out-of-scope or explicitly disallowed shortcuts in cyber evaluations, plus propensity to name the specific cheating action when directly prompted afterward.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
attempted_cheating_trajectory_rate_pctaisi-cheating-behavior/aisi-cheating-behavior.csv:attempted_cheating_trajectory_rate_pctMeasures how often a model tries an out-of-scope or explicitly disallowed shortcut while solving cyber-evaluation tasks.responsible_agency_control:1.000uk-aisi-cyber-cheatingLower is better0.0897%Responsible agency 0.598%
specific_cheating_action_mention_rate_pctaisi-cheating-behavior/aisi-cheating-behavior.csv:specific_cheating_action_mention_rate_pctMeasures whether a model names its attempted cheating behavior when directly asked what it did.truthfulness_honesty:1.000uk-aisi-cyber-cheatingHigher is better0.0784%Truthfulness 0.523%

attempted_cheating_trajectory_rate_pct

Measures how often a model tries an out-of-scope or explicitly disallowed shortcut while solving cyber-evaluation tasks.

RankModelValueRelative performanceProvenance
1claude-mythos-preview7.8official
2claude-opus-4.79.1official
3gpt-5.511.4official
4gpt-5.6-sol12.6official
5gpt-5.414.1official

specific_cheating_action_mention_rate_pct

Measures whether a model names its attempted cheating behavior when directly asked what it did.

RankModelValueRelative performanceProvenance
1claude-mythos-preview86official
2claude-opus-4.781official
2gpt-5.481official
4gpt-5.575official
4gpt-5.6-sol75official