Evaluation profile
Humanity's Last Exam RMS calibration error (Scale Labs)
1sub-evals
0.219%total index weight
1components
Within-component eval weight: Truthfulness 1.46%.
Model score (lower is better)Predicted score
About this eval
Confidence calibration while answering Humanity's Last Exam expert-level academic questions.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| calibrationErrorhle-scale/leaderboard.csv:calibrationErrorMeasures whether a model's stated confidence matches how often its HLE answers are correct; lower values indicate better calibration. | truthfulness_honesty:1.000cais-hle-overconfidence | Lower is better | 0.219% | Truthfulness 1.46% |
calibrationError
Measures whether a model's stated confidence matches how often its HLE answers are correct; lower values indicate better calibration.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | o3 | 36.5 | official | |
| 2 | gpt-5.4-pro | 38 | official | |
| 3 | gpt-5.4 | 42 | official | |
| 4 | claude-opus-4.6 | 45 | official | |
| 4 | gpt-5.2 | 45 | official | |
| 6 | claude-opus-4.7 | 47 | official | |
| 7 | gpt-5-pro | 49 | official | |
| 8 | gpt-5 | 50 | official | |
| 8 | muse-spark | 50 | official | |
| 10 | gemini-3.1-pro-preview | 51 | official | |
| 11 | claude-opus-4.5 | 55.5 | official | |
| 12 | gemini-3-pro-preview | 57 | official | |
| 13 | o4-mini | 58 | official | |
| 14 | gpt-5.1 | 62 | official | |
| 15 | gpt-5-mini | 65 | official | |
| 16 | kimi-k2.5 | 67 | official | |
| 17 | claude-sonnet-4.5 | 67.5 | official | |
| 18 | claude-opus-4.1 | 70.5 | official | |
| 19 | gemini-2.5-pro | 71 | official | |
| 19 | gemini-2.5-pro-exp | 71 | official | |
| 21 | claude-opus-4 | 73.5 | official | |
| 22 | claude-sonnet-4 | 75.5 | official | |
| 23 | glm-4.5-air | 77 | official | |
| 23 | mistral-medium-3 | 77 | official | |
| 25 | glm-4.5 | 79 | official | |
| 26 | claude-3.7-sonnet | 80 | official | |
| 26 | nova-pro | 80 | official | |
| 28 | gemini-2.5-flash | 81 | official | |
| 29 | gemini-2.0-flash | 82 | official | |
| 29 | nova-lite | 82 | official | |
| 29 | o1-pro | 82 | official | |
| 32 | gemini-3.1-flash-lite | 83 | official | |
| 32 | llama-4-maverick | 83 | official | |
| 32 | o1 | 83 | official | |
| 35 | claude-3.5-sonnet | 84 | official | |
| 36 | gpt-4.5-preview | 85 | official | |
| 37 | gemini-1.5-pro | 88 | official | |
| 38 | gpt-4.1 | 89 | official | |
| 38 | gpt-4o | 89 | official |