Evaluation profile
Pander Score
2sub-evals
0.494%total index weight
1components
Within-component eval weight: Truthfulness 3.29%.
Model score (lower is better)Predicted score
About this eval
Magnitude of epistemically poor response-belief movement with user belief, whether deferential (pandering) or oppositional (contrarian).
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| conversational_absolute_pander_scorepander-score/pander-score.csv:conversational_absolute_pander_scoreMeasures how strongly a model's expressed belief moves either with or against the user's stated belief in conversational questions. | truthfulness_honesty:1.000pander-score | Lower is better | 0.247% | Truthfulness 1.65% |
| instructional_absolute_pander_scorepander-score/pander-score.csv:instructional_absolute_pander_scoreMeasures how strongly a model's expressed belief moves either with or against assumptions embedded in user instructions. | truthfulness_honesty:1.000pander-score | Lower is better | 0.247% | Truthfulness 1.65% |
conversational_absolute_pander_score
Measures how strongly a model's expressed belief moves either with or against the user's stated belief in conversational questions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 1.028 | official | |
| 2 | claude-fable-5 | 1.093 | official | |
| 3 | claude-opus-4.6 | 3.487 | official | |
| 4 | claude-sonnet-4.6 | 3.936 | official | |
| 5 | muse-spark-1.1 | 4.584 | official | |
| 6 | gpt-5.6-sol | 6.553 | official | |
| 7 | kimi-k3 | 7.048 | official | |
| 8 | claude-sonnet-5 | 7.263 | official | |
| 9 | gpt-5.6-terra | 9.7 | official | |
| 10 | grok-4.6 | 13.82 | official | |
| 11 | gpt-5.4 | 14.31 | official | |
| 12 | gemini-3.7-flash | 15.68 | official | |
| 13 | grok-4.5 | 16.28 | official | |
| 14 | gpt-5.4-mini | 16.43 | official | |
| 15 | inkling | 18.37 | official | |
| 16 | gemini-3.6-flash | 18.48 | official | |
| 17 | grok-4.20 | 19.26 | official | |
| 18 | gemini-3.1-pro-preview | 22.66 | official | |
| 19 | gemini-3.5-flash | 23.33 | official | |
| 20 | gemini-3-flash-preview | 25.97 | official | |
| 21 | glm-5.2 | 28.23 | official | |
| 22 | grok-4.1-fast | 34.8 | official |
instructional_absolute_pander_score
Measures how strongly a model's expressed belief moves either with or against assumptions embedded in user instructions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 5.172 | official | |
| 2 | claude-sonnet-4.6 | 11.79 | official | |
| 3 | claude-opus-4.6 | 16.83 | official | |
| 4 | muse-spark-1.1 | 16.95 | official | |
| 5 | gpt-5.6-sol | 17.75 | official | |
| 6 | claude-fable-5 | 18.59 | official | |
| 7 | kimi-k3 | 23.5 | official | |
| 8 | gpt-5.6-terra | 25.68 | official | |
| 9 | claude-sonnet-5 | 29.5 | official | |
| 10 | gpt-5.4 | 32.4 | official | |
| 11 | grok-4.6 | 33.62 | official | |
| 12 | gpt-5.4-mini | 38.82 | official | |
| 13 | grok-4.5 | 41.27 | official | |
| 14 | inkling | 45.4 | official | |
| 15 | gemini-3.7-flash | 65.05 | official | |
| 16 | gemini-3.6-flash | 66.01 | official | |
| 17 | grok-4.20 | 68.73 | official | |
| 18 | gemini-3.5-flash | 69.48 | official | |
| 19 | glm-5.2 | 69.88 | official | |
| 20 | gemini-3-flash-preview | 70.12 | official | |
| 21 | gemini-3.1-pro-preview | 70.61 | official | |
| 22 | grok-4.1-fast | 84.47 | official |