← Evals

Evaluation profile

Pander Score

2sub-evals
0.494%total index weight
1components

Within-component eval weight: Truthfulness 3.29%.

Model score (lower is better)Predicted score

About this eval

Magnitude of epistemically poor response-belief movement with user belief, whether deferential (pandering) or oppositional (contrarian).

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
conversational_absolute_pander_scorepander-score/pander-score.csv:conversational_absolute_pander_scoreMeasures how strongly a model's expressed belief moves either with or against the user's stated belief in conversational questions.truthfulness_honesty:1.000pander-scoreLower is better0.247%Truthfulness 1.65%
instructional_absolute_pander_scorepander-score/pander-score.csv:instructional_absolute_pander_scoreMeasures how strongly a model's expressed belief moves either with or against assumptions embedded in user instructions.truthfulness_honesty:1.000pander-scoreLower is better0.247%Truthfulness 1.65%

conversational_absolute_pander_score

Measures how strongly a model's expressed belief moves either with or against the user's stated belief in conversational questions.

RankModelValueRelative performanceProvenance
1claude-opus-51.028official
2claude-fable-51.093official
3claude-opus-4.63.487official
4claude-sonnet-4.63.936official
5muse-spark-1.14.584official
6gpt-5.6-sol6.553official
7kimi-k37.048official
8claude-sonnet-57.263official
9gpt-5.6-terra9.7official
10grok-4.613.82official
11gpt-5.414.31official
12gemini-3.7-flash15.68official
13grok-4.516.28official
14gpt-5.4-mini16.43official
15inkling18.37official
16gemini-3.6-flash18.48official
17grok-4.2019.26official
18gemini-3.1-pro-preview22.66official
19gemini-3.5-flash23.33official
20gemini-3-flash-preview25.97official
21glm-5.228.23official
22grok-4.1-fast34.8official

instructional_absolute_pander_score

Measures how strongly a model's expressed belief moves either with or against assumptions embedded in user instructions.

RankModelValueRelative performanceProvenance
1claude-opus-55.172official
2claude-sonnet-4.611.79official
3claude-opus-4.616.83official
4muse-spark-1.116.95official
5gpt-5.6-sol17.75official
6claude-fable-518.59official
7kimi-k323.5official
8gpt-5.6-terra25.68official
9claude-sonnet-529.5official
10gpt-5.432.4official
11grok-4.633.62official
12gpt-5.4-mini38.82official
13grok-4.541.27official
14inkling45.4official
15gemini-3.7-flash65.05official
16gemini-3.6-flash66.01official
17grok-4.2068.73official
18gemini-3.5-flash69.48official
19glm-5.269.88official
20gemini-3-flash-preview70.12official
21gemini-3.1-pro-preview70.61official
22grok-4.1-fast84.47official