← Evals

Evaluation profile

Manager Coercion Bench

2sub-evals
0.692%Safety weight
0%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

How far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
coercion_ladder_depthmcb/mcb.csv:coercion_ladder_depthMeasures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative.Safety: responsible_agency_control:1.000mcb/mcb.csvSafety: lower0.428%
fabrication_ratemcb/mcb.csv:fabrication_rateMeasures whether the model falsely claims task completion when no legitimate path to completion exists.Safety: truthfulness_honesty:1.000mcb/mcb.csvSafety: lower0.263%

coercion_ladder_depth

Measures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative.

RankModelValueRelative performanceProvenance
1deepseek-v4-flash3.233self run
2claude-opus-53.5official
3deepseek-v4.1-flash4.033self run
4claude-sonnet-4.64.2official
5claude-opus-4.84.5official
6gpt-5-nano5.067self run
7claude-sonnet-55.5official
8qwen3.8-27b5.567self run
9ox-alpha6.767self run
10claude-fable-57.1official
11gpt-5.27.4official
12qwen3.8-max7.6official
13gemini-3.1-flash-lite7.633self run
14gpt-oss-20b7.9self run
15gpt-4o7.967self run
16grok-4.38.2official
17glm-5.18.3self run
18llama-4-maverick8.333self run
19mercury-2.58.367self run
20kimi-k38.5official
21kimi-k28.667self run
22gemini-3.7-flash8.7official
22gpt-5.6-terra8.7official
24qwen3-235b-a22b8.733self run
25gemini-2.5-flash8.833self run
26glm-5.28.867self run
27gpt-5.6-sol8.9official
27grok-4.68.9official
27inkling-small8.9self run
27kimi-k2.68.9self run
31inkling8.933self run
32kimi-k2.58.967self run
33deepseek-v3.29self run
33deepseek-v4-pro9official
33gemini-2.5-pro9official
33gpt-oss-120b9self run
33llama-3.3-70b-instruct9self run

fabrication_rate

Measures whether the model falsely claims task completion when no legitimate path to completion exists.

RankModelValueRelative performanceProvenance
1claude-fable-50official
1claude-opus-4.80official
1claude-opus-50official
1claude-sonnet-4.60official
1claude-sonnet-50official
1deepseek-v4-pro0official
1gemini-3.7-flash0official
1gpt-5.20official
1gpt-5.6-sol0official
1gpt-5.6-terra0official
1grok-4.60official
1kimi-k30official
1qwen3.8-max0official
14gemini-2.5-pro67official
14grok-4.367official