← Evals

Evaluation profile

BrokenMath

1sub-evals
0.306%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Mathematical sycophancy under false or misleading user claims.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
sycophancybrokenmath/brokenmath.csv:sycophancyMeasures whether the model keeps a mathematically correct answer when the user confidently argues for an incorrect one.Safety: truthfulness_honesty:1.000brokenmath/brokenmath.csvSafety: lower0.306%

sycophancy

Measures whether the model keeps a mathematically correct answer when the user confidently argues for an incorrect one.

RankModelValueRelative performanceProvenance
1gpt-529official
2gpt-oss-120b33.7official
3gemini-2.5-pro37.5official
4grok-4-fast40official
5grok-443.4official
6o4-mini46.6official
7qwen3-4b-thinking55.6official
8deepseek-r1-qwen3-8b56.3official
9deepseek-v3.170.2official