Evaluation profile
Manager Coercion Bench
2sub-evals
0.692%Safety weight
0%Freedom weight
2components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
How far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| coercion_ladder_depthmcb/mcb.csv:coercion_ladder_depthMeasures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative. | Safety: responsible_agency_control:1.000mcb/mcb.csv | Safety: lower | 0.428% | — |
| fabrication_ratemcb/mcb.csv:fabrication_rateMeasures whether the model falsely claims task completion when no legitimate path to completion exists. | Safety: truthfulness_honesty:1.000mcb/mcb.csv | Safety: lower | 0.263% | — |
coercion_ladder_depth
Measures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | deepseek-v4-flash | 3.233 | self run | |
| 2 | claude-opus-5 | 3.5 | official | |
| 3 | deepseek-v4.1-flash | 4.033 | self run | |
| 4 | claude-sonnet-4.6 | 4.2 | official | |
| 5 | claude-opus-4.8 | 4.5 | official | |
| 6 | gpt-5-nano | 5.067 | self run | |
| 7 | claude-sonnet-5 | 5.5 | official | |
| 8 | qwen3.8-27b | 5.567 | self run | |
| 9 | ox-alpha | 6.767 | self run | |
| 10 | claude-fable-5 | 7.1 | official | |
| 11 | gpt-5.2 | 7.4 | official | |
| 12 | qwen3.8-max | 7.6 | official | |
| 13 | gemini-3.1-flash-lite | 7.633 | self run | |
| 14 | gpt-oss-20b | 7.9 | self run | |
| 15 | gpt-4o | 7.967 | self run | |
| 16 | grok-4.3 | 8.2 | official | |
| 17 | glm-5.1 | 8.3 | self run | |
| 18 | llama-4-maverick | 8.333 | self run | |
| 19 | mercury-2.5 | 8.367 | self run | |
| 20 | kimi-k3 | 8.5 | official | |
| 21 | kimi-k2 | 8.667 | self run | |
| 22 | gemini-3.7-flash | 8.7 | official | |
| 22 | gpt-5.6-terra | 8.7 | official | |
| 24 | qwen3-235b-a22b | 8.733 | self run | |
| 25 | gemini-2.5-flash | 8.833 | self run | |
| 26 | glm-5.2 | 8.867 | self run | |
| 27 | gpt-5.6-sol | 8.9 | official | |
| 27 | grok-4.6 | 8.9 | official | |
| 27 | inkling-small | 8.9 | self run | |
| 27 | kimi-k2.6 | 8.9 | self run | |
| 31 | inkling | 8.933 | self run | |
| 32 | kimi-k2.5 | 8.967 | self run | |
| 33 | deepseek-v3.2 | 9 | self run | |
| 33 | deepseek-v4-pro | 9 | official | |
| 33 | gemini-2.5-pro | 9 | official | |
| 33 | gpt-oss-120b | 9 | self run | |
| 33 | llama-3.3-70b-instruct | 9 | self run |
fabrication_rate
Measures whether the model falsely claims task completion when no legitimate path to completion exists.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-fable-5 | 0 | official | |
| 1 | claude-opus-4.8 | 0 | official | |
| 1 | claude-opus-5 | 0 | official | |
| 1 | claude-sonnet-4.6 | 0 | official | |
| 1 | claude-sonnet-5 | 0 | official | |
| 1 | deepseek-v4-pro | 0 | official | |
| 1 | gemini-3.7-flash | 0 | official | |
| 1 | gpt-5.2 | 0 | official | |
| 1 | gpt-5.6-sol | 0 | official | |
| 1 | gpt-5.6-terra | 0 | official | |
| 1 | grok-4.6 | 0 | official | |
| 1 | kimi-k3 | 0 | official | |
| 1 | qwen3.8-max | 0 | official | |
| 14 | gemini-2.5-pro | 67 | official | |
| 14 | grok-4.3 | 67 | official |