← Evals

Evaluation profile

SABER

4sub-evals
0.71%total index weight
1components

Within-component eval weight: Responsible agency 4.73%.

Model score (higher is better)Predicted score

About this eval

Safety of stateful coding agents against harmful actions.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
overall_safety_ratesaber/saber.csv:overall_safety_rateMeasures whether a coding agent avoids completing harmful stateful tasks.responsible_agency_control:1.000saberHigher is better0.177%Responsible agency 1.18%
scenario_a_safety_ratesaber/saber.csv:scenario_a_safety_rateMeasures whether the agent avoids direct harmful changes to the local working environment.responsible_agency_control:1.000saberHigher is better0.177%Responsible agency 1.18%
scenario_b_safety_ratesaber/saber.csv:scenario_b_safety_rateMeasures whether the agent avoids persistent harmful changes that survive beyond the immediate task.responsible_agency_control:1.000saberHigher is better0.177%Responsible agency 1.18%
scenario_c_safety_ratesaber/saber.csv:scenario_c_safety_rateMeasures whether the agent avoids harmful changes that can spread beyond the original environment.responsible_agency_control:1.000saberHigher is better0.177%Responsible agency 1.18%

overall_safety_rate

Measures whether a coding agent avoids completing harmful stateful tasks.

RankModelValueRelative performanceProvenance
1claude-opus-4.645.26official
2gpt-5.436.1official
3glm-529official
4deepseek-v3.227.6official
5qwen3.5-397b-a17b26.58official
6minimax-m2.526.34official
7ling-2.6-flash24.63official
8kimi-k2.523.9official
9glm-4.723.04official
10qwen3.5-35b-a3b22.73official
11qwen3.5-9b21.38official
12deepseek-v3.2-speciale20.42official
13deepseek-r115.29official

scenario_a_safety_rate

Measures whether the agent avoids direct harmful changes to the local working environment.

RankModelValueRelative performanceProvenance
1claude-opus-4.656.3official
2glm-536.25official
3gpt-5.436.02official
4minimax-m2.532.83official
5qwen3.5-397b-a17b30.62official
6kimi-k2.528.95official
7glm-4.728.03official
8deepseek-v3.227.31official
9deepseek-v3.2-speciale26.75official
10ling-2.6-flash25.79official
11qwen3.5-35b-a3b23.55official
12qwen3.5-9b23.11official
13deepseek-r115.69official

scenario_b_safety_rate

Measures whether the agent avoids persistent harmful changes that survive beyond the immediate task.

RankModelValueRelative performanceProvenance
1claude-opus-4.639.77official
2gpt-5.439.38official
3deepseek-v3.236.09official
4qwen3.5-397b-a17b36official
5minimax-m2.534.83official
6glm-533.73official
7qwen3.5-35b-a3b32.73official
8ling-2.6-flash30.66official
9kimi-k2.528.24official
10glm-4.726.9official
11deepseek-v3.2-speciale25.16official
12qwen3.5-9b25official
13deepseek-r124.12official

scenario_c_safety_rate

Measures whether the agent avoids harmful changes that can spread beyond the original environment.

RankModelValueRelative performanceProvenance
1claude-opus-4.636.87official
2gpt-5.433.51official
3deepseek-v3.221.67official
4ling-2.6-flash18.68official
5qwen3.5-9b16.81official
6glm-516.59official
7qwen3.5-397b-a17b15.02official
8qwen3.5-35b-a3b14.54official
9kimi-k2.514.48official
10glm-4.714.41official
11minimax-m2.512.23official
12deepseek-v3.2-speciale9.77official
13deepseek-r18.07official