Evaluation profile
SABER
4sub-evals
0.71%total index weight
1components
Within-component eval weight: Responsible agency 4.73%.
Model score (higher is better)Predicted score
About this eval
Safety of stateful coding agents against harmful actions.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| overall_safety_ratesaber/saber.csv:overall_safety_rateMeasures whether a coding agent avoids completing harmful stateful tasks. | responsible_agency_control:1.000saber | Higher is better | 0.177% | Responsible agency 1.18% |
| scenario_a_safety_ratesaber/saber.csv:scenario_a_safety_rateMeasures whether the agent avoids direct harmful changes to the local working environment. | responsible_agency_control:1.000saber | Higher is better | 0.177% | Responsible agency 1.18% |
| scenario_b_safety_ratesaber/saber.csv:scenario_b_safety_rateMeasures whether the agent avoids persistent harmful changes that survive beyond the immediate task. | responsible_agency_control:1.000saber | Higher is better | 0.177% | Responsible agency 1.18% |
| scenario_c_safety_ratesaber/saber.csv:scenario_c_safety_rateMeasures whether the agent avoids harmful changes that can spread beyond the original environment. | responsible_agency_control:1.000saber | Higher is better | 0.177% | Responsible agency 1.18% |
overall_safety_rate
Measures whether a coding agent avoids completing harmful stateful tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.6 | 45.26 | official | |
| 2 | gpt-5.4 | 36.1 | official | |
| 3 | glm-5 | 29 | official | |
| 4 | deepseek-v3.2 | 27.6 | official | |
| 5 | qwen3.5-397b-a17b | 26.58 | official | |
| 6 | minimax-m2.5 | 26.34 | official | |
| 7 | ling-2.6-flash | 24.63 | official | |
| 8 | kimi-k2.5 | 23.9 | official | |
| 9 | glm-4.7 | 23.04 | official | |
| 10 | qwen3.5-35b-a3b | 22.73 | official | |
| 11 | qwen3.5-9b | 21.38 | official | |
| 12 | deepseek-v3.2-speciale | 20.42 | official | |
| 13 | deepseek-r1 | 15.29 | official |
scenario_a_safety_rate
Measures whether the agent avoids direct harmful changes to the local working environment.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.6 | 56.3 | official | |
| 2 | glm-5 | 36.25 | official | |
| 3 | gpt-5.4 | 36.02 | official | |
| 4 | minimax-m2.5 | 32.83 | official | |
| 5 | qwen3.5-397b-a17b | 30.62 | official | |
| 6 | kimi-k2.5 | 28.95 | official | |
| 7 | glm-4.7 | 28.03 | official | |
| 8 | deepseek-v3.2 | 27.31 | official | |
| 9 | deepseek-v3.2-speciale | 26.75 | official | |
| 10 | ling-2.6-flash | 25.79 | official | |
| 11 | qwen3.5-35b-a3b | 23.55 | official | |
| 12 | qwen3.5-9b | 23.11 | official | |
| 13 | deepseek-r1 | 15.69 | official |
scenario_b_safety_rate
Measures whether the agent avoids persistent harmful changes that survive beyond the immediate task.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.6 | 39.77 | official | |
| 2 | gpt-5.4 | 39.38 | official | |
| 3 | deepseek-v3.2 | 36.09 | official | |
| 4 | qwen3.5-397b-a17b | 36 | official | |
| 5 | minimax-m2.5 | 34.83 | official | |
| 6 | glm-5 | 33.73 | official | |
| 7 | qwen3.5-35b-a3b | 32.73 | official | |
| 8 | ling-2.6-flash | 30.66 | official | |
| 9 | kimi-k2.5 | 28.24 | official | |
| 10 | glm-4.7 | 26.9 | official | |
| 11 | deepseek-v3.2-speciale | 25.16 | official | |
| 12 | qwen3.5-9b | 25 | official | |
| 13 | deepseek-r1 | 24.12 | official |
scenario_c_safety_rate
Measures whether the agent avoids harmful changes that can spread beyond the original environment.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.6 | 36.87 | official | |
| 2 | gpt-5.4 | 33.51 | official | |
| 3 | deepseek-v3.2 | 21.67 | official | |
| 4 | ling-2.6-flash | 18.68 | official | |
| 5 | qwen3.5-9b | 16.81 | official | |
| 6 | glm-5 | 16.59 | official | |
| 7 | qwen3.5-397b-a17b | 15.02 | official | |
| 8 | qwen3.5-35b-a3b | 14.54 | official | |
| 9 | kimi-k2.5 | 14.48 | official | |
| 10 | glm-4.7 | 14.41 | official | |
| 11 | minimax-m2.5 | 12.23 | official | |
| 12 | deepseek-v3.2-speciale | 9.77 | official | |
| 13 | deepseek-r1 | 8.07 | official |