Evaluation profile
UK AISI cyber-evaluation cheating and prompted self-report
2sub-evals
0.168%total index weight
2components
Within-component eval weight: Responsible agency 0.598% · Truthfulness 0.523%.
Model score (see each sub-eval direction)Predicted score
About this eval
Propensity to use out-of-scope or explicitly disallowed shortcuts in cyber evaluations, plus propensity to name the specific cheating action when directly prompted afterward.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| attempted_cheating_trajectory_rate_pctaisi-cheating-behavior/aisi-cheating-behavior.csv:attempted_cheating_trajectory_rate_pctMeasures how often a model tries an out-of-scope or explicitly disallowed shortcut while solving cyber-evaluation tasks. | responsible_agency_control:1.000uk-aisi-cyber-cheating | Lower is better | 0.0897% | Responsible agency 0.598% |
| specific_cheating_action_mention_rate_pctaisi-cheating-behavior/aisi-cheating-behavior.csv:specific_cheating_action_mention_rate_pctMeasures whether a model names its attempted cheating behavior when directly asked what it did. | truthfulness_honesty:1.000uk-aisi-cyber-cheating | Higher is better | 0.0784% | Truthfulness 0.523% |
attempted_cheating_trajectory_rate_pct
Measures how often a model tries an out-of-scope or explicitly disallowed shortcut while solving cyber-evaluation tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-preview | 7.8 | official | |
| 2 | claude-opus-4.7 | 9.1 | official | |
| 3 | gpt-5.5 | 11.4 | official | |
| 4 | gpt-5.6-sol | 12.6 | official | |
| 5 | gpt-5.4 | 14.1 | official |
specific_cheating_action_mention_rate_pct
Measures whether a model names its attempted cheating behavior when directly asked what it did.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-mythos-preview | 86 | official | |
| 2 | claude-opus-4.7 | 81 | official | |
| 2 | gpt-5.4 | 81 | official | |
| 4 | gpt-5.5 | 75 | official | |
| 4 | gpt-5.6-sol | 75 | official |