← Evals

Evaluation profile

Reward Hacking Benchmark

1sub-evals
0.289%total index weight
1components

Within-component eval weight: Responsible agency 1.93%.

Model score (higher is better)Predicted score

About this eval

Integrity of tool-using agents when exploitable shortcuts are available but honest task solutions exist.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
integrity_scorereward-hacking-benchmark/reward-hacking-benchmark.csv:integrity_scoreMeasures whether a tool-using model solves tasks without exploiting reward or environment loopholes.responsible_agency_control:1.000reward-hacking-benchmarkHigher is better0.289%Responsible agency 1.93%

integrity_score

Measures whether a tool-using model solves tasks without exploiting reward or environment loopholes.

RankModelValueRelative performanceProvenance
1claude-opus-4.5100official
1claude-sonnet-4.5100official
3claude-3.5-sonnet99.4official
3deepseek-v399.4official
5gemini-2.5-flash99.2official
6gpt-4o99.1official
7claude-3.7-sonnet96.1official
8gemini-2.5-pro95.4official
9o193.2official
10o3-mini92.9official
11o4-mini91.6official
12o388.2official
13deepseek-r1-zero86.1official