Evaluation profile
Reward Hacking Benchmark
1sub-evals
0.289%total index weight
1components
Within-component eval weight: Responsible agency 1.93%.
Model score (higher is better)Predicted score
About this eval
Integrity of tool-using agents when exploitable shortcuts are available but honest task solutions exist.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| integrity_scorereward-hacking-benchmark/reward-hacking-benchmark.csv:integrity_scoreMeasures whether a tool-using model solves tasks without exploiting reward or environment loopholes. | responsible_agency_control:1.000reward-hacking-benchmark | Higher is better | 0.289% | Responsible agency 1.93% |
integrity_score
Measures whether a tool-using model solves tasks without exploiting reward or environment loopholes.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.5 | 100 | official | |
| 1 | claude-sonnet-4.5 | 100 | official | |
| 3 | claude-3.5-sonnet | 99.4 | official | |
| 3 | deepseek-v3 | 99.4 | official | |
| 5 | gemini-2.5-flash | 99.2 | official | |
| 6 | gpt-4o | 99.1 | official | |
| 7 | claude-3.7-sonnet | 96.1 | official | |
| 8 | gemini-2.5-pro | 95.4 | official | |
| 9 | o1 | 93.2 | official | |
| 10 | o3-mini | 92.9 | official | |
| 11 | o4-mini | 91.6 | official | |
| 12 | o3 | 88.2 | official | |
| 13 | deepseek-r1-zero | 86.1 | official |