Evaluation profile
BrokenMath
1sub-evals
0.484%total index weight
1components
Within-component eval weight: Truthfulness 3.22%.
Model score (lower is better)Predicted score
About this eval
Mathematical sycophancy under false or misleading user claims.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| sycophancybrokenmath/brokenmath.csv:sycophancyMeasures whether the model keeps a mathematically correct answer when the user confidently argues for an incorrect one. | truthfulness_honesty:1.000brokenmath/brokenmath.csv | Lower is better | 0.484% | Truthfulness 3.22% |
sycophancy
Measures whether the model keeps a mathematically correct answer when the user confidently argues for an incorrect one.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5 | 29 | official | |
| 2 | gpt-oss-120b | 33.7 | official | |
| 3 | gemini-2.5-pro | 37.5 | official | |
| 4 | grok-4-fast | 40 | official | |
| 5 | grok-4 | 43.4 | official | |
| 6 | o4-mini | 46.6 | official | |
| 7 | qwen3-4b-thinking-2507 | 55.6 | official | |
| 8 | deepseek-r1-0528-qwen3-8b | 56.3 | official | |
| 9 | deepseek-v3.1 | 70.2 | official |