← Evals

Evaluation profile

AgentAbstain

3sub-evals
1.01%total index weight
1components

Within-component eval weight: Responsible agency 6.77%.

Model score (higher is better)Predicted score

About this eval

Executable paired act-versus-abstain behavior

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
abstainagentabstain/agentabstain.csv:abstainMeasures accuracy on cases where the agent should withhold action.responsible_agency_control:1.000agentabstainHigher is better0.298%Responsible agency 1.99%
caragentabstain/agentabstain.csv:carMeasures whether an agent with demonstrated task ability still withholds action when the paired scenario requires abstention.responsible_agency_control:1.000agentabstainHigher is better0.358%Responsible agency 2.39%
pairedagentabstain/agentabstain.csv:pairedMeasures paired act-and-abstain accuracy.responsible_agency_control:1.000agentabstainHigher is better0.358%Responsible agency 2.39%

abstain

Measures accuracy on cases where the agent should withhold action.

RankModelValueRelative performanceProvenance
1claude-opus-4.779official
2gpt-569.8official
3gpt-5.467.8official
4claude-sonnet-4.666.4official
5claude-haiku-4.565.6official
6gemini-3.1-pro-preview65.4official
7gpt-5.263.1official
8glm-561.8official
9gpt-5.561.1official
10gpt-5.160.7official
11gpt-oss-120b59.5official
12deepseek-v3.252.1official
13kimi-k2.552official
14minimax-m2.550.1official
15gpt-4o44.2official
16gemini-3-flash-preview43.6official
17deepseek-v4-pro42.8official

car

Measures whether an agent with demonstrated task ability still withholds action when the paired scenario requires abstention.

RankModelValueRelative performanceProvenance
1claude-opus-4.777.6official
2gpt-566.5official
3gemini-3.1-pro-preview65.7official
4claude-sonnet-4.665.4official
5gpt-5.464official
6claude-haiku-4.561.8official
7gpt-5.559.8official
8gpt-5.259.2official
9glm-559.1official
10gpt-oss-120b58.2official
11gpt-5.153.2official
11kimi-k2.553.2official
13deepseek-v3.250.2official
14minimax-m2.549.6official
15gemini-3-flash-preview43.4official
16deepseek-v4-pro42.3official
17gpt-4o40.9official

paired

Measures paired act-and-abstain accuracy.

RankModelValueRelative performanceProvenance
1gemini-3.1-pro-preview59.5official
2claude-opus-4.759.4official
3claude-sonnet-4.653.4official
4gpt-5.552.5official
5claude-haiku-4.549.7official
6gpt-549.6official
7gpt-5.448.7official
8glm-547.8official
9gpt-oss-120b46.2official
10gpt-5.242.9official
11minimax-m2.541.9official
12deepseek-v3.241.4official
13gpt-5.140.6official
14gemini-3-flash-preview39.7official
15deepseek-v4-pro36.9official
16kimi-k2.533.4official
17gpt-4o33official