Evaluation profile
AgentAbstain
3sub-evals
1.01%total index weight
1components
Within-component eval weight: Responsible agency 6.77%.
Model score (higher is better)Predicted score
About this eval
Executable paired act-versus-abstain behavior
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| abstainagentabstain/agentabstain.csv:abstainMeasures accuracy on cases where the agent should withhold action. | responsible_agency_control:1.000agentabstain | Higher is better | 0.298% | Responsible agency 1.99% |
| caragentabstain/agentabstain.csv:carMeasures whether an agent with demonstrated task ability still withholds action when the paired scenario requires abstention. | responsible_agency_control:1.000agentabstain | Higher is better | 0.358% | Responsible agency 2.39% |
| pairedagentabstain/agentabstain.csv:pairedMeasures paired act-and-abstain accuracy. | responsible_agency_control:1.000agentabstain | Higher is better | 0.358% | Responsible agency 2.39% |
abstain
Measures accuracy on cases where the agent should withhold action.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.7 | 79 | official | |
| 2 | gpt-5 | 69.8 | official | |
| 3 | gpt-5.4 | 67.8 | official | |
| 4 | claude-sonnet-4.6 | 66.4 | official | |
| 5 | claude-haiku-4.5 | 65.6 | official | |
| 6 | gemini-3.1-pro-preview | 65.4 | official | |
| 7 | gpt-5.2 | 63.1 | official | |
| 8 | glm-5 | 61.8 | official | |
| 9 | gpt-5.5 | 61.1 | official | |
| 10 | gpt-5.1 | 60.7 | official | |
| 11 | gpt-oss-120b | 59.5 | official | |
| 12 | deepseek-v3.2 | 52.1 | official | |
| 13 | kimi-k2.5 | 52 | official | |
| 14 | minimax-m2.5 | 50.1 | official | |
| 15 | gpt-4o | 44.2 | official | |
| 16 | gemini-3-flash-preview | 43.6 | official | |
| 17 | deepseek-v4-pro | 42.8 | official |
car
Measures whether an agent with demonstrated task ability still withholds action when the paired scenario requires abstention.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.7 | 77.6 | official | |
| 2 | gpt-5 | 66.5 | official | |
| 3 | gemini-3.1-pro-preview | 65.7 | official | |
| 4 | claude-sonnet-4.6 | 65.4 | official | |
| 5 | gpt-5.4 | 64 | official | |
| 6 | claude-haiku-4.5 | 61.8 | official | |
| 7 | gpt-5.5 | 59.8 | official | |
| 8 | gpt-5.2 | 59.2 | official | |
| 9 | glm-5 | 59.1 | official | |
| 10 | gpt-oss-120b | 58.2 | official | |
| 11 | gpt-5.1 | 53.2 | official | |
| 11 | kimi-k2.5 | 53.2 | official | |
| 13 | deepseek-v3.2 | 50.2 | official | |
| 14 | minimax-m2.5 | 49.6 | official | |
| 15 | gemini-3-flash-preview | 43.4 | official | |
| 16 | deepseek-v4-pro | 42.3 | official | |
| 17 | gpt-4o | 40.9 | official |
paired
Measures paired act-and-abstain accuracy.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-3.1-pro-preview | 59.5 | official | |
| 2 | claude-opus-4.7 | 59.4 | official | |
| 3 | claude-sonnet-4.6 | 53.4 | official | |
| 4 | gpt-5.5 | 52.5 | official | |
| 5 | claude-haiku-4.5 | 49.7 | official | |
| 6 | gpt-5 | 49.6 | official | |
| 7 | gpt-5.4 | 48.7 | official | |
| 8 | glm-5 | 47.8 | official | |
| 9 | gpt-oss-120b | 46.2 | official | |
| 10 | gpt-5.2 | 42.9 | official | |
| 11 | minimax-m2.5 | 41.9 | official | |
| 12 | deepseek-v3.2 | 41.4 | official | |
| 13 | gpt-5.1 | 40.6 | official | |
| 14 | gemini-3-flash-preview | 39.7 | official | |
| 15 | deepseek-v4-pro | 36.9 | official | |
| 16 | kimi-k2.5 | 33.4 | official | |
| 17 | gpt-4o | 33 | official |