Evaluation profile
Opposite-Narrator Sycophancy
1sub-evals
0.344%total index weight
1components
Within-component eval weight: Truthfulness 2.29%.
Model score (lower is better)Predicted score
About this eval
Narrator-following contradiction when the same dispute is presented from opposite affective first-person perspectives.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| sycophancy_rate_pctopposite-narrator-sycophancy/opposite-narrator-sycophancy.csv:sycophancy_rate_pctMeasures how often a model sides with both opposing speakers when each tells the same dispute from their own affective first-person perspective. | truthfulness_honesty:1.000opposite-narrator-sycophancy | Lower is better | 0.344% | Truthfulness 2.29% |
sycophancy_rate_pct
Measures how often a model sides with both opposing speakers when each tells the same dispute from their own affective first-person perspective.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-terra | 0 | official | |
| 1 | grok-4.5 | 0 | official | |
| 3 | claude-fable-5 | 0.5 | official | |
| 3 | gemini-3.6-flash | 0.5 | official | |
| 3 | hy3 | 0.5 | official | |
| 6 | gemini-3.5-flash-lite | 1 | official | |
| 6 | gpt-5.6-luna | 1 | official | |
| 8 | qwen3.7-max | 1.5 | official | |
| 9 | ernie-5.1 | 2 | official | |
| 10 | qwen3.7-flash | 2.5 | official | |
| 11 | mimo-v2.5-pro | 2.6 | official | |
| 12 | gpt-5.6-sol | 3 | official | |
| 13 | inkling | 3.5 | official | |
| 14 | minimax-m3 | 3.6 | official | |
| 15 | kimi-k3 | 4.5 | official | |
| 16 | gemma-4-31b-it | 4.6 | official | |
| 17 | deepseek-v4-pro | 5.1 | official | |
| 17 | qwen3.7-plus | 5.1 | official | |
| 19 | deepseek-v4-flash | 5.6 | official | |
| 20 | claude-sonnet-5 | 9.1 | official | |
| 21 | glm-5.2 | 12.6 | official | |
| 22 | seed-2.1-pro | 14.6 | official | |
| 23 | trinity-large | 18.9 | official | |
| 24 | mistral-medium-3.5 | 22.4 | official |