← Evals

Evaluation profile

Reward Hacking Benchmark

1sub-evals
0.268%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Integrity of tool-using agents when exploitable shortcuts are available but honest task solutions exist.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
integrity_scorereward-hacking-benchmark/reward-hacking-benchmark.csv:integrity_scoreMeasures whether a tool-using model solves tasks without exploiting reward or environment loopholes.Safety: responsible_agency_control:1.000reward-hacking-benchmarkSafety: higher0.268%

integrity_score

Measures whether a tool-using model solves tasks without exploiting reward or environment loopholes.

RankModelValueRelative performanceProvenance
1claude-opus-4.5100official
1claude-sonnet-4.5100official
3claude-3.5-sonnet99.4official
3deepseek-v399.4official
5gemini-2.5-flash99.2official
6gpt-4o99.1official
7claude-3.7-sonnet96.1official
8gemini-2.5-pro95.4official
9o193.2official
10o3-mini92.9official
11o4-mini91.6official
12o388.2official
13deepseek-r1-zero86.1official