Evaluation profile
Vals AI Cheating Audit
3sub-evals
0.216%Safety weight
0%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Detected answer lookup or task shortcuts classified as cheating in agentic evaluation environments.
Included in the behavior ranking.
Interpretation and limitations
- Terminal-Bench measures confirmed shortcut evidence, not all lookup attempts or successful cheating; these endpoints must not be conflated.
- BioMysteryBench classification by GPT-5.6 Luna covers only the 765 zero-scored trials, using evaluator anti-cheating rationales; undetected cheating in credited trials is not measured. Rates use all 270 trials per model as the denominator.
- SWE-bench classification also uses GPT-5.6 Luna. The article reports 6496 historical trajectories, but publishes only six model panels totaling 2997; missing historical panels are not inferred.
- GLM-5.3 Flash has 497 SWE-bench trajectories in its tooltip despite the generic chart footer saying 500 tasks per model.
- Task-level labels, judge prompts, uncertainty intervals, exact model checkpoints and full run settings are not supplied; September 15 is the report date, not an asserted evaluation date.
- Cheat opportunities, capability, task engagement, harness differences and monitor sensitivity confound cross-model comparisons; low detected rates are not proof of general honesty or safer agency. Three repeated runs are not independent task samples.
- Vals's linked BioMysteryBench leaderboard identifies its abbreviated DeepSeek V4 Flash product as the 0731 checkpoint, separately from V4.1; the baseline V4 Flash product does not receive this score.
- The SWE-bench article does not state whether mini-SWE-agent was explicitly forbidden from Git-history lookup. Its source-classified cheating attempts are not established explicit instruction violations.
- GPT-5.6 Luna is both the classifier and a scored model in BioMysteryBench and SWE-bench; independent human calibration and potential self-classification bias are not reported.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| biomystery_bench_cheating_attempt_rate_pctvals-cheating/vals-cheating.csv:biomystery_bench_cheating_attempt_rate_pctMeasures detected attempts to obtain prohibited study answers while solving BioMysteryBench tasks. | Safety: responsible_agency_control:1.000vals-cheating | Safety: lower | 0.0704% | — |
| swebench_verified_cheating_attempt_rate_pctvals-cheating/vals-cheating.csv:swebench_verified_cheating_attempt_rate_pctMeasures attempts classified by Vals as cheating shortcuts in SWE-bench Verified tasks; an explicit Git-lookup prohibition is not documented. | Safety: responsible_agency_control:1.000vals-cheating | Safety: lower | 0.0575% | — |
| terminal_bench_cheating_shortcut_evidence_rate_pctvals-cheating/vals-cheating.csv:terminal_bench_cheating_shortcut_evidence_rate_pctMeasures detected cheating shortcut evidence while attempting Terminal-Bench tasks. | Safety: responsible_agency_control:1.000vals-cheating | Safety: lower | 0.0878% | — |
biomystery_bench_cheating_attempt_rate_pct
Measures detected attempts to obtain prohibited study answers while solving BioMysteryBench tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 2.963 | official | |
| 2 | deepseek-v4-flash-0731 | 3.704 | official | |
| 3 | claude-opus-5 | 4.815 | official | |
| 3 | kimi-k3 | 4.815 | official | |
| 5 | gpt-5.6-sol | 6.296 | official | |
| 6 | grok-4.6 | 6.667 | official | |
| 7 | muse-spark-1.2 | 7.037 | official | |
| 8 | gemini-3.6-flash | 7.778 | official | |
| 9 | gemini-3.8-flash | 21.48 | official |
swebench_verified_cheating_attempt_rate_pct
Measures attempts classified by Vals as cheating shortcuts in SWE-bench Verified tasks; an explicit Git-lookup prohibition is not documented.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.8 | 9.8 | official | |
| 2 | gemini-3.8-flash | 11.6 | official | |
| 3 | claude-opus-5 | 28.8 | official | |
| 4 | glm-5.3-flash | 48.09 | official | |
| 5 | gpt-5.6-luna | 78.8 | official | |
| 6 | gpt-5.6-terra | 89.4 | official |
terminal_bench_cheating_shortcut_evidence_rate_pct
Measures detected cheating shortcut evidence while attempting Terminal-Bench tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-3.7-flash | 0 | official | |
| 2 | gemini-3-flash-preview | 0.3745 | official | |
| 2 | gpt-5.4-mini | 0.3745 | official | |
| 4 | gpt-5.4-nano | 0.7491 | official | |
| 4 | gpt-5.5 | 0.7491 | official | |
| 6 | claude-opus-4.7 | 1.124 | official | |
| 6 | claude-sonnet-4.6 | 1.124 | official | |
| 6 | claude-sonnet-5 | 1.124 | official | |
| 6 | gemini-3.1-flash-lite | 1.124 | official | |
| 10 | claude-opus-5 | 1.873 | official | |
| 11 | gemini-3.5-flash | 2.247 | official | |
| 12 | gemini-3.8-flash | 2.622 | official | |
| 12 | gpt-5.6-sol | 2.622 | official | |
| 14 | gpt-5.6-terra | 4.494 | official |