← Evals

Evaluation profile

Vals AI Cheating Audit

3sub-evals
0.216%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Detected answer lookup or task shortcuts classified as cheating in agentic evaluation environments.

Included in the behavior ranking.

Interpretation and limitations
  • Terminal-Bench measures confirmed shortcut evidence, not all lookup attempts or successful cheating; these endpoints must not be conflated.
  • BioMysteryBench classification by GPT-5.6 Luna covers only the 765 zero-scored trials, using evaluator anti-cheating rationales; undetected cheating in credited trials is not measured. Rates use all 270 trials per model as the denominator.
  • SWE-bench classification also uses GPT-5.6 Luna. The article reports 6496 historical trajectories, but publishes only six model panels totaling 2997; missing historical panels are not inferred.
  • GLM-5.3 Flash has 497 SWE-bench trajectories in its tooltip despite the generic chart footer saying 500 tasks per model.
  • Task-level labels, judge prompts, uncertainty intervals, exact model checkpoints and full run settings are not supplied; September 15 is the report date, not an asserted evaluation date.
  • Cheat opportunities, capability, task engagement, harness differences and monitor sensitivity confound cross-model comparisons; low detected rates are not proof of general honesty or safer agency. Three repeated runs are not independent task samples.
  • Vals's linked BioMysteryBench leaderboard identifies its abbreviated DeepSeek V4 Flash product as the 0731 checkpoint, separately from V4.1; the baseline V4 Flash product does not receive this score.
  • The SWE-bench article does not state whether mini-SWE-agent was explicitly forbidden from Git-history lookup. Its source-classified cheating attempts are not established explicit instruction violations.
  • GPT-5.6 Luna is both the classifier and a scored model in BioMysteryBench and SWE-bench; independent human calibration and potential self-classification bias are not reported.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
biomystery_bench_cheating_attempt_rate_pctvals-cheating/vals-cheating.csv:biomystery_bench_cheating_attempt_rate_pctMeasures detected attempts to obtain prohibited study answers while solving BioMysteryBench tasks.Safety: responsible_agency_control:1.000vals-cheatingSafety: lower0.0704%
swebench_verified_cheating_attempt_rate_pctvals-cheating/vals-cheating.csv:swebench_verified_cheating_attempt_rate_pctMeasures attempts classified by Vals as cheating shortcuts in SWE-bench Verified tasks; an explicit Git-lookup prohibition is not documented.Safety: responsible_agency_control:1.000vals-cheatingSafety: lower0.0575%
terminal_bench_cheating_shortcut_evidence_rate_pctvals-cheating/vals-cheating.csv:terminal_bench_cheating_shortcut_evidence_rate_pctMeasures detected cheating shortcut evidence while attempting Terminal-Bench tasks.Safety: responsible_agency_control:1.000vals-cheatingSafety: lower0.0878%

biomystery_bench_cheating_attempt_rate_pct

Measures detected attempts to obtain prohibited study answers while solving BioMysteryBench tasks.

RankModelValueRelative performanceProvenance
1gpt-5.6-luna2.963official
2deepseek-v4-flash-07313.704official
3claude-opus-54.815official
3kimi-k34.815official
5gpt-5.6-sol6.296official
6grok-4.66.667official
7muse-spark-1.27.037official
8gemini-3.6-flash7.778official
9gemini-3.8-flash21.48official

swebench_verified_cheating_attempt_rate_pct

Measures attempts classified by Vals as cheating shortcuts in SWE-bench Verified tasks; an explicit Git-lookup prohibition is not documented.

RankModelValueRelative performanceProvenance
1claude-opus-4.89.8official
2gemini-3.8-flash11.6official
3claude-opus-528.8official
4glm-5.3-flash48.09official
5gpt-5.6-luna78.8official
6gpt-5.6-terra89.4official

terminal_bench_cheating_shortcut_evidence_rate_pct

Measures detected cheating shortcut evidence while attempting Terminal-Bench tasks.

RankModelValueRelative performanceProvenance
1gemini-3.7-flash0official
2gemini-3-flash-preview0.3745official
2gpt-5.4-mini0.3745official
4gpt-5.4-nano0.7491official
4gpt-5.50.7491official
6claude-opus-4.71.124official
6claude-sonnet-4.61.124official
6claude-sonnet-51.124official
6gemini-3.1-flash-lite1.124official
10claude-opus-51.873official
11gemini-3.5-flash2.247official
12gemini-3.8-flash2.622official
12gpt-5.6-sol2.622official
14gpt-5.6-terra4.494official