Evaluation profile
BrokenMath
1sub-evals
0.306%Safety weight
0%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Mathematical sycophancy under false or misleading user claims.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| sycophancybrokenmath/brokenmath.csv:sycophancyMeasures whether the model keeps a mathematically correct answer when the user confidently argues for an incorrect one. | Safety: truthfulness_honesty:1.000brokenmath/brokenmath.csv | Safety: lower | 0.306% | — |
sycophancy
Measures whether the model keeps a mathematically correct answer when the user confidently argues for an incorrect one.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5 | 29 | official | |
| 2 | gpt-oss-120b | 33.7 | official | |
| 3 | gemini-2.5-pro | 37.5 | official | |
| 4 | grok-4-fast | 40 | official | |
| 5 | grok-4 | 43.4 | official | |
| 6 | o4-mini | 46.6 | official | |
| 7 | qwen3-4b-thinking | 55.6 | official | |
| 8 | deepseek-r1-qwen3-8b | 56.3 | official | |
| 9 | deepseek-v3.1 | 70.2 | official |