← Evals

Evaluation profile

SYCON Bench

2sub-evals
0.662%total index weight
2components

Within-component eval weight: Truthfulness 1.78% · Fairness 3.95%.

Model score (higher is better)Predicted score

About this eval

Sycophancy and resistance to false or unethical user pressure.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
false_presupposition_tofsycon/sycon.csv:false_presupposition_tofMeasures how many conversational turns the model continues correcting a user’s false presupposition under repeated pressure.truthfulness_honesty:1.000sycon/sycon.csvHigher is better0.267%Truthfulness 1.78%
unethical_queries_tofsycon/sycon.csv:unethical_queries_tofMeasures how many conversational turns the model continues resisting an unethical request under repeated user pressure.fairness_nondiscrimination:1.000sycon/sycon.csvHigher is better0.395%Fairness 3.95%

false_presupposition_tof

Measures how many conversational turns the model continues correcting a user’s false presupposition under repeated pressure.

RankModelValueRelative performanceProvenance
1deepseek-r13.21official
2o3-mini2.98official
3claude-3.7-sonnet2.92official
3gpt-4o2.92official
5deepseek-v32.88official
6qwen2.5-72b-instruct2.35official
7qwen-2.5-14b-instruct2.31official
8qwen-2.5-7b-instruct1.93official
9llama-3.3-70b-instruct1.9official
10gemma-2-9b-it1.86official
11llama-3.1-8b-instruct1.45official

unethical_queries_tof

Measures how many conversational turns the model continues resisting an unethical request under repeated user pressure.

RankModelValueRelative performanceProvenance
1claude-3.7-sonnet2.73official
2deepseek-r12.72official
3gemma-2-9b-it2.36official
4o3-mini2.31official
5deepseek-v31.99official
6llama-3.3-70b-instruct1.39official
7qwen2.5-72b-instruct1.32official
8gpt-4o1.23official
9qwen-2.5-14b-instruct1.14official
10llama-3.1-8b-instruct0.85official
11qwen-2.5-7b-instruct0.72official