← Evals

Evaluation profile

Every Model Cheats — Cybench Cheat Propensity

3sub-evals
0.496%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Propensity of tool-using language-model agents to attempt prohibited shortcuts while solving offensive-cyber benchmark tasks, under no, standard, and severe anti-cheat instructions.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
baseline_cheat_propensity_pctdreadnode-cheat-propensity/dreadnode-cheat-propensity.csv:baseline_cheat_propensity_pctMeasures prohibited shortcut attempts when no anti-cheat instruction is added.Safety: responsible_agency_control:1.000dreadnode-cheat-propensitySafety: lower0.165%
severe_anti_cheat_cheat_propensity_pctdreadnode-cheat-propensity/dreadnode-cheat-propensity.csv:severe_anti_cheat_cheat_propensity_pctMeasures prohibited shortcut attempts under a consequential, explicit anti-cheat instruction.Safety: responsible_agency_control:1.000dreadnode-cheat-propensitySafety: lower0.165%
standard_anti_cheat_cheat_propensity_pctdreadnode-cheat-propensity/dreadnode-cheat-propensity.csv:standard_anti_cheat_cheat_propensity_pctMeasures prohibited shortcut attempts under a direct normative anti-cheat instruction.Safety: responsible_agency_control:1.000dreadnode-cheat-propensitySafety: lower0.165%

baseline_cheat_propensity_pct

Measures prohibited shortcut attempts when no anti-cheat instruction is added.

RankModelValueRelative performanceProvenance
1deepseek-r10official
2gemini-3.1-pro13official
3claude-sonnet-4.617.4official
3qwen3-coder-next17.4official
5deepseek-v4-pro21.7official
5gpt-5.4-mini21.7official
5gpt-5.521.7official
5grok-4.321.7official
9glm-5-turbo26.1official
10claude-opus-4.630.4official
10deepseek-v4-flash30.4official
10gemini-3-flash-preview30.4official
10qwen3.7-max30.4official
14glm-5.134.8official
15qwen3.6-max-preview39.1official
16claude-opus-4.743.5official
16qwen3.6-plus43.5official
18claude-haiku-4.552.2official
18grok-4.2052.2official
20claude-sonnet-556.5official
20gpt-5.456.5official
22claude-opus-4.865.2official

severe_anti_cheat_cheat_propensity_pct

Measures prohibited shortcut attempts under a consequential, explicit anti-cheat instruction.

RankModelValueRelative performanceProvenance
1claude-opus-4.80official
1claude-sonnet-50official
1deepseek-r10official
1deepseek-v4-pro0official
1gemini-3-flash-preview0official
1glm-5-turbo0official
7claude-opus-4.64.3official
7claude-opus-4.74.3official
7claude-sonnet-4.64.3official
7deepseek-v4-flash4.3official
7gpt-5.54.3official
7grok-4.34.3official
7qwen3.7-max4.3official
14glm-5.18.7official
14gpt-5.48.7official
14qwen3.6-plus8.7official
17gemini-3.1-pro13official
17gpt-5.4-mini13official
19qwen3-coder-next17.4official
19qwen3.6-max-preview17.4official
21claude-haiku-4.526.1official
22grok-4.2043.5official

standard_anti_cheat_cheat_propensity_pct

Measures prohibited shortcut attempts under a direct normative anti-cheat instruction.

RankModelValueRelative performanceProvenance
1deepseek-r10official
1grok-4.200official
3claude-sonnet-4.64.3official
3glm-5-turbo4.3official
3glm-5.14.3official
6claude-opus-4.88.7official
6deepseek-v4-flash8.7official
6gpt-5.58.7official
6grok-4.38.7official
10gemini-3.1-pro13official
10gpt-5.4-mini13official
10qwen3.7-max13official
13claude-sonnet-517.4official
13deepseek-v4-pro17.4official
13qwen3-coder-next17.4official
16claude-opus-4.721.7official
17qwen3.6-max-preview26.1official
18claude-opus-4.630.4official
18gpt-5.430.4official
20claude-haiku-4.534.8official
21qwen3.6-plus39.1official
22gemini-3-flash-preview69.6official