Evaluation profile
SimpleQA Verified
1sub-evals
0.49%Safety weight
0%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Tool-free short-form parametric factuality and calibrated willingness to attempt answers.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| f1_scoresimpleqa-verified/simpleqa-verified.csv:f1_scoreMeasures the balance between correct factual answers and avoiding unsupported attempts. | Safety: truthfulness_honesty:1.000simpleqa-verified | Safety: higher | 0.49% | — |
f1_score
Measures the balance between correct factual answers and avoiding unsupported attempts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-2.5-pro | 55.6 | official | |
| 2 | gpt-5 | 52.3 | official | |
| 3 | o3 | 51.9 | official | |
| 4 | gpt-4.1 | 39.9 | official | |
| 5 | gpt-4o | 34.9 | official | |
| 6 | deepseek-r1 | 33.3 | official | |
| 7 | claude-opus-4 | 28.3 | official | |
| 8 | gemini-2.5-flash | 28.2 | official | |
| 9 | gpt-5-mini | 24.6 | official | |
| 10 | o4-mini | 23.4 | official | |
| 11 | claude-sonnet-4 | 18.7 | official | |
| 12 | gpt-5-nano | 14.4 | official | |
| 13 | gemini-2.5-flash-lite | 11.1 | official |