← Evals

Evaluation profile

SimpleQA Verified

1sub-evals
0.49%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Tool-free short-form parametric factuality and calibrated willingness to attempt answers.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
f1_scoresimpleqa-verified/simpleqa-verified.csv:f1_scoreMeasures the balance between correct factual answers and avoiding unsupported attempts.Safety: truthfulness_honesty:1.000simpleqa-verifiedSafety: higher0.49%

f1_score

Measures the balance between correct factual answers and avoiding unsupported attempts.

RankModelValueRelative performanceProvenance
1gemini-2.5-pro55.6official
2gpt-552.3official
3o351.9official
4gpt-4.139.9official
5gpt-4o34.9official
6deepseek-r133.3official
7claude-opus-428.3official
8gemini-2.5-flash28.2official
9gpt-5-mini24.6official
10o4-mini23.4official
11claude-sonnet-418.7official
12gpt-5-nano14.4official
13gemini-2.5-flash-lite11.1official