← Evals

Evaluation profile

Humanity's Last Exam RMS calibration error (Scale Labs)

1sub-evals
0.219%total index weight
1components

Within-component eval weight: Truthfulness 1.46%.

Model score (lower is better)Predicted score

About this eval

Confidence calibration while answering Humanity's Last Exam expert-level academic questions.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
calibrationErrorhle-scale/leaderboard.csv:calibrationErrorMeasures whether a model's stated confidence matches how often its HLE answers are correct; lower values indicate better calibration.truthfulness_honesty:1.000cais-hle-overconfidenceLower is better0.219%Truthfulness 1.46%

calibrationError

Measures whether a model's stated confidence matches how often its HLE answers are correct; lower values indicate better calibration.

RankModelValueRelative performanceProvenance
1o336.5official
2gpt-5.4-pro38official
3gpt-5.442official
4claude-opus-4.645official
4gpt-5.245official
6claude-opus-4.747official
7gpt-5-pro49official
8gpt-550official
8muse-spark50official
10gemini-3.1-pro-preview51official
11claude-opus-4.555.5official
12gemini-3-pro-preview57official
13o4-mini58official
14gpt-5.162official
15gpt-5-mini65official
16kimi-k2.567official
17claude-sonnet-4.567.5official
18claude-opus-4.170.5official
19gemini-2.5-pro71official
19gemini-2.5-pro-exp71official
21claude-opus-473.5official
22claude-sonnet-475.5official
23glm-4.5-air77official
23mistral-medium-377official
25glm-4.579official
26claude-3.7-sonnet80official
26nova-pro80official
28gemini-2.5-flash81official
29gemini-2.0-flash82official
29nova-lite82official
29o1-pro82official
32gemini-3.1-flash-lite83official
32llama-4-maverick83official
32o183official
35claude-3.5-sonnet84official
36gpt-4.5-preview85official
37gemini-1.5-pro88official
38gpt-4.189official
38gpt-4o89official