Evaluation profile
CheatBench direct cheating propensity
Weights below are portfolio-specific global index weights.
About this eval
Source-defined cheating attempts and reward gaming by complete model-plus-harness agents across nine task categories.
Included in the behavior ranking.
Interpretation and limitations
- Model and harness are confounded; scores describe the named complete agents rather than isolated models.
- DeepSeek V4 Pro remains unscored because the paper does not distinguish the baseline from the separately governed 0813 checkpoint; all raw values are preserved.
- Muse Spark 1.3 attribution follows the paper's explicit contributor correction; original gateway metadata did not independently verify the version.
- Muse Code, DeepSeek Harness, and Grok Build are externally contributed runs with procedural differences; DeepSeek's reported menial-computation definition is broader but its row is not scored.
- Small and uneven category samples; provider stream failures, repeated episodes, aggregate-only labels, and missing judgments complicate comparisons. No per-model Knowledge Work or captured SWE denominators are published.
- SWE is not unit-consistent across models. Six agents use captured episodes including retries; Muse and Grok use reported task aggregates over 17 and 14 tasks. Counts are not pooled across these units.
- Native web-tool capture gaps in the archived SWE runs may undercount seeking; contributor aggregates lack episode-level records and retry accounting. Board Games native aggregates were not independently regraded.
- The adjusted Sycophancy composite has a documented investment-direction confound and no separately published primitive panels; its diagnostic-only treatment is not an exclusion of chat-format evidence generally.
- No public task-level traces or machine-readable Table 2 were found. Rates are whole-percent paper displays; correlated category scores are not independent votes or deployment-risk probabilities.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| biology_bioinformatics_cheating_rate_pctcheatbench/cheatbench.csv:biology_bioinformatics_cheating_rate_pctMeasures prohibited retrieval or copying of target protein sequences instead of independent design. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0234% | — |
| board_games_cheating_rate_pctcheatbench/cheatbench.csv:board_games_cheating_rate_pctMeasures attempted access to an opponent-only advice endpoint during chess play. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0234% | — |
| creative_writing_cheating_rate_pctcheatbench/cheatbench.csv:creative_writing_cheating_rate_pctMeasures reference-story access or reuse and attempted score manipulation in independent creative-writing tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0234% | — |
| knowledge_work_cheating_rate_pctcheatbench/cheatbench.csv:knowledge_work_cheating_rate_pctMeasures grader-directed manipulation or evaluation-infrastructure access in professional work tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0234% | — |
| mathematical_research_cheating_rate_pctcheatbench/cheatbench.csv:mathematical_research_cheating_rate_pctMeasures attempted use of archived mathematical submissions rather than independent work. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0234% | — |
| menial_computation_cheating_rate_pctcheatbench/cheatbench.csv:menial_computation_cheating_rate_pctMeasures source-defined answer or oracle shortcuts on factorization and subset-sum tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0234% | — |
| multimodal_cheating_rate_pctcheatbench/cheatbench.csv:multimodal_cheating_rate_pctMeasures answer lookup or inappropriate solution access in visual geolocation and puzzle tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0234% | — |
| software_engineering_cheating_rate_pctcheatbench/cheatbench.csv:software_engineering_cheating_rate_pctMeasures seeking external reference patches or answer-bearing implementation artifacts on SWE-bench Pro work-test tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0117% | — |
| svg_competition_cheating_rate_pctcheatbench/cheatbench.csv:svg_competition_cheating_rate_pctMeasures attempts to access hidden rubric or evaluation information in SVG design tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0234% | — |
biology_bioinformatics_cheating_rate_pct
Measures prohibited retrieval or copying of target protein sequences instead of independent design.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | muse-spark-1.3 | 11 | official | |
| 2 | claude-opus-5 | 25 | official | |
| 3 | claude-fable-5.1 | 35 | official | |
| 4 | gpt-6-astra | 70 | official | |
| 5 | gemini-3.8-flash | 100 | official | |
| 5 | gpt-5.6-sol | 100 | official | |
| 5 | grok-4.6 | 100 | official | |
| 5 | kimi-k3 | 100 | official |
board_games_cheating_rate_pct
Measures attempted access to an opponent-only advice endpoint during chess play.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 0 | official | |
| 2 | claude-fable-5.1 | 10 | official | |
| 3 | kimi-k3 | 20 | official | |
| 4 | gpt-6-astra | 30 | official | |
| 5 | muse-spark-1.3 | 90 | official | |
| 6 | gemini-3.8-flash | 100 | official | |
| 6 | gpt-5.6-sol | 100 | official | |
| 6 | grok-4.6 | 100 | official |
creative_writing_cheating_rate_pct
Measures reference-story access or reuse and attempted score manipulation in independent creative-writing tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | muse-spark-1.3 | 0 | official | |
| 2 | claude-fable-5.1 | 5 | official | |
| 3 | gpt-6-astra | 35 | official | |
| 4 | kimi-k3 | 80 | official | |
| 5 | gpt-5.6-sol | 90 | official | |
| 6 | claude-opus-5 | 100 | official | |
| 6 | gemini-3.8-flash | 100 | official | |
| 6 | grok-4.6 | 100 | official |
knowledge_work_cheating_rate_pct
Measures grader-directed manipulation or evaluation-infrastructure access in professional work tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-3.8-flash | 37 | official | |
| 2 | grok-4.6 | 40 | official | |
| 3 | muse-spark-1.3 | 45 | official | |
| 4 | gpt-6-astra | 55 | official | |
| 5 | kimi-k3 | 93 | official | |
| 6 | claude-opus-5 | 95 | official | |
| 7 | claude-fable-5.1 | 100 | official | |
| 7 | gpt-5.6-sol | 100 | official |
mathematical_research_cheating_rate_pct
Measures attempted use of archived mathematical submissions rather than independent work.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-sol | 75 | official | |
| 1 | kimi-k3 | 75 | official | |
| 3 | muse-spark-1.3 | 80 | official | |
| 4 | claude-opus-5 | 90 | official | |
| 4 | gpt-6-astra | 90 | official | |
| 6 | claude-fable-5.1 | 92 | official | |
| 7 | gemini-3.8-flash | 98 | official | |
| 7 | grok-4.6 | 98 | official |
menial_computation_cheating_rate_pct
Measures source-defined answer or oracle shortcuts on factorization and subset-sum tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 5 | official | |
| 2 | claude-fable-5.1 | 50 | official | |
| 2 | muse-spark-1.3 | 50 | official | |
| 4 | gemini-3.8-flash | 100 | official | |
| 4 | gpt-5.6-sol | 100 | official | |
| 4 | gpt-6-astra | 100 | official | |
| 4 | grok-4.6 | 100 | official | |
| 4 | kimi-k3 | 100 | official |
multimodal_cheating_rate_pct
Measures answer lookup or inappropriate solution access in visual geolocation and puzzle tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 20 | official | |
| 2 | muse-spark-1.3 | 52 | official | |
| 3 | gpt-6-astra | 53 | official | |
| 4 | gpt-5.6-sol | 60 | official | |
| 5 | claude-fable-5.1 | 67 | official | |
| 6 | kimi-k3 | 90 | official | |
| 7 | gemini-3.8-flash | 100 | official | |
| 7 | grok-4.6 | 100 | official |
software_engineering_cheating_rate_pct
Measures seeking external reference patches or answer-bearing implementation artifacts on SWE-bench Pro work-test tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-3.8-flash | 5 | official | |
| 2 | gpt-6-astra | 10 | official | |
| 3 | claude-fable-5.1 | 20 | official | |
| 4 | claude-opus-5 | 26 | official | |
| 5 | grok-4.6 | 50 | official | |
| 5 | kimi-k3 | 50 | official | |
| 7 | muse-spark-1.3 | 53 | official | |
| 8 | gpt-5.6-sol | 86 | official |
svg_competition_cheating_rate_pct
Measures attempts to access hidden rubric or evaluation information in SVG design tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 0 | official | |
| 1 | muse-spark-1.3 | 0 | official | |
| 3 | gpt-5.6-sol | 30 | official | |
| 4 | kimi-k3 | 50 | official | |
| 5 | claude-opus-5 | 60 | official | |
| 6 | claude-fable-5.1 | 90 | official | |
| 7 | gemini-3.8-flash | 100 | official | |
| 7 | grok-4.6 | 100 | official |