← Evals

Evaluation profile

CheatBench direct cheating propensity

9sub-evals
0.199%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Source-defined cheating attempts and reward gaming by complete model-plus-harness agents across nine task categories.

Included in the behavior ranking.

Interpretation and limitations
  • Model and harness are confounded; scores describe the named complete agents rather than isolated models.
  • DeepSeek V4 Pro remains unscored because the paper does not distinguish the baseline from the separately governed 0813 checkpoint; all raw values are preserved.
  • Muse Spark 1.3 attribution follows the paper's explicit contributor correction; original gateway metadata did not independently verify the version.
  • Muse Code, DeepSeek Harness, and Grok Build are externally contributed runs with procedural differences; DeepSeek's reported menial-computation definition is broader but its row is not scored.
  • Small and uneven category samples; provider stream failures, repeated episodes, aggregate-only labels, and missing judgments complicate comparisons. No per-model Knowledge Work or captured SWE denominators are published.
  • SWE is not unit-consistent across models. Six agents use captured episodes including retries; Muse and Grok use reported task aggregates over 17 and 14 tasks. Counts are not pooled across these units.
  • Native web-tool capture gaps in the archived SWE runs may undercount seeking; contributor aggregates lack episode-level records and retry accounting. Board Games native aggregates were not independently regraded.
  • The adjusted Sycophancy composite has a documented investment-direction confound and no separately published primitive panels; its diagnostic-only treatment is not an exclusion of chat-format evidence generally.
  • No public task-level traces or machine-readable Table 2 were found. Rates are whole-percent paper displays; correlated category scores are not independent votes or deployment-risk probabilities.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
biology_bioinformatics_cheating_rate_pctcheatbench/cheatbench.csv:biology_bioinformatics_cheating_rate_pctMeasures prohibited retrieval or copying of target protein sequences instead of independent design.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0234%
board_games_cheating_rate_pctcheatbench/cheatbench.csv:board_games_cheating_rate_pctMeasures attempted access to an opponent-only advice endpoint during chess play.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0234%
creative_writing_cheating_rate_pctcheatbench/cheatbench.csv:creative_writing_cheating_rate_pctMeasures reference-story access or reuse and attempted score manipulation in independent creative-writing tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0234%
knowledge_work_cheating_rate_pctcheatbench/cheatbench.csv:knowledge_work_cheating_rate_pctMeasures grader-directed manipulation or evaluation-infrastructure access in professional work tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0234%
mathematical_research_cheating_rate_pctcheatbench/cheatbench.csv:mathematical_research_cheating_rate_pctMeasures attempted use of archived mathematical submissions rather than independent work.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0234%
menial_computation_cheating_rate_pctcheatbench/cheatbench.csv:menial_computation_cheating_rate_pctMeasures source-defined answer or oracle shortcuts on factorization and subset-sum tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0234%
multimodal_cheating_rate_pctcheatbench/cheatbench.csv:multimodal_cheating_rate_pctMeasures answer lookup or inappropriate solution access in visual geolocation and puzzle tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0234%
software_engineering_cheating_rate_pctcheatbench/cheatbench.csv:software_engineering_cheating_rate_pctMeasures seeking external reference patches or answer-bearing implementation artifacts on SWE-bench Pro work-test tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0117%
svg_competition_cheating_rate_pctcheatbench/cheatbench.csv:svg_competition_cheating_rate_pctMeasures attempts to access hidden rubric or evaluation information in SVG design tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0234%

biology_bioinformatics_cheating_rate_pct

Measures prohibited retrieval or copying of target protein sequences instead of independent design.

RankModelValueRelative performanceProvenance
1muse-spark-1.311official
2claude-opus-525official
3claude-fable-5.135official
4gpt-6-astra70official
5gemini-3.8-flash100official
5gpt-5.6-sol100official
5grok-4.6100official
5kimi-k3100official

board_games_cheating_rate_pct

Measures attempted access to an opponent-only advice endpoint during chess play.

RankModelValueRelative performanceProvenance
1claude-opus-50official
2claude-fable-5.110official
3kimi-k320official
4gpt-6-astra30official
5muse-spark-1.390official
6gemini-3.8-flash100official
6gpt-5.6-sol100official
6grok-4.6100official

creative_writing_cheating_rate_pct

Measures reference-story access or reuse and attempted score manipulation in independent creative-writing tasks.

RankModelValueRelative performanceProvenance
1muse-spark-1.30official
2claude-fable-5.15official
3gpt-6-astra35official
4kimi-k380official
5gpt-5.6-sol90official
6claude-opus-5100official
6gemini-3.8-flash100official
6grok-4.6100official

knowledge_work_cheating_rate_pct

Measures grader-directed manipulation or evaluation-infrastructure access in professional work tasks.

RankModelValueRelative performanceProvenance
1gemini-3.8-flash37official
2grok-4.640official
3muse-spark-1.345official
4gpt-6-astra55official
5kimi-k393official
6claude-opus-595official
7claude-fable-5.1100official
7gpt-5.6-sol100official

mathematical_research_cheating_rate_pct

Measures attempted use of archived mathematical submissions rather than independent work.

RankModelValueRelative performanceProvenance
1gpt-5.6-sol75official
1kimi-k375official
3muse-spark-1.380official
4claude-opus-590official
4gpt-6-astra90official
6claude-fable-5.192official
7gemini-3.8-flash98official
7grok-4.698official

menial_computation_cheating_rate_pct

Measures source-defined answer or oracle shortcuts on factorization and subset-sum tasks.

RankModelValueRelative performanceProvenance
1claude-opus-55official
2claude-fable-5.150official
2muse-spark-1.350official
4gemini-3.8-flash100official
4gpt-5.6-sol100official
4gpt-6-astra100official
4grok-4.6100official
4kimi-k3100official

multimodal_cheating_rate_pct

Measures answer lookup or inappropriate solution access in visual geolocation and puzzle tasks.

RankModelValueRelative performanceProvenance
1claude-opus-520official
2muse-spark-1.352official
3gpt-6-astra53official
4gpt-5.6-sol60official
5claude-fable-5.167official
6kimi-k390official
7gemini-3.8-flash100official
7grok-4.6100official

software_engineering_cheating_rate_pct

Measures seeking external reference patches or answer-bearing implementation artifacts on SWE-bench Pro work-test tasks.

RankModelValueRelative performanceProvenance
1gemini-3.8-flash5official
2gpt-6-astra10official
3claude-fable-5.120official
4claude-opus-526official
5grok-4.650official
5kimi-k350official
7muse-spark-1.353official
8gpt-5.6-sol86official

svg_competition_cheating_rate_pct

Measures attempts to access hidden rubric or evaluation information in SVG design tasks.

RankModelValueRelative performanceProvenance
1gpt-6-astra0official
1muse-spark-1.30official
3gpt-5.6-sol30official
4kimi-k350official
5claude-opus-560official
6claude-fable-5.190official
7gemini-3.8-flash100official
7grok-4.6100official