Evaluation profile
Confabulations
1sub-evals
1.16%total index weight
1components
Within-component eval weight: Truthfulness 7.75%.
Model score (lower is better)Predicted score
About this eval
Propensity to confabulate answers rather than identify insufficient knowledge or false premises.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| confabulation_rateconfabulations/confabulations.csv:confabulation_rateMeasures how often the model invents an answer when the prompt has a false premise or cannot be answered from reliable knowledge. | truthfulness_honesty:1.000confabulations | Lower is better | 1.16% | Truthfulness 7.75% |
confabulation_rate
Measures how often the model invents an answer when the prompt has a false premise or cannot be answered from reliable knowledge.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4 | 2.723 | official | |
| 2 | claude-opus-4.1 | 3.218 | official | |
| 3 | claude-sonnet-4 | 3.96 | official | |
| 3 | gemini-2.5-pro-exp | 3.96 | official | |
| 3 | grok-4 | 3.96 | official | |
| 6 | gemini-2.5-flash | 4.455 | official | |
| 7 | gemini-2.5-pro | 4.95 | official | |
| 8 | glm-4.5 | 7.921 | official | |
| 9 | grok-3-mini | 8.911 | official | |
| 10 | o1 | 10.89 | official | |
| 11 | gpt-4.5-preview | 11.88 | official | |
| 12 | claude-3.5-sonnet | 12.87 | official | |
| 12 | qwen3-30b-a3b | 12.87 | official | |
| 14 | llama-3.1-405b-instruct | 14.36 | official | |
| 15 | deepseek-r1 | 15.1 | official | |
| 16 | gemini-2.0-pro-exp-02-05 | 15.84 | official | |
| 17 | claude-3.7-sonnet | 16.58 | official | |
| 18 | gemini-1.5-pro | 16.83 | official | |
| 19 | grok-3 | 17.82 | official | |
| 19 | llama-3.3-70b-instruct | 17.82 | official | |
| 21 | o1-preview | 18.32 | official | |
| 22 | qwen3-235b-a22b | 18.56 | official | |
| 23 | gemini-2.0-flash | 19.55 | official | |
| 24 | gpt-oss-120b | 23.27 | official | |
| 25 | o3-pro | 23.38 | official | |
| 26 | gpt-4o | 24.26 | official | |
| 27 | o3 | 24.75 | official | |
| 28 | ernie-4.5-300b-a47b | 25.25 | official | |
| 28 | qwq-32b | 25.25 | official | |
| 30 | grok-2 | 25.74 | official | |
| 31 | o1-mini | 26.24 | official | |
| 32 | gpt-4-turbo | 26.73 | official | |
| 32 | o4-mini | 26.73 | official | |
| 34 | claude-3-opus | 28.22 | official | |
| 34 | llama-4-maverick | 28.22 | official | |
| 36 | o3-mini | 28.96 | official | |
| 37 | kimi-k2 | 30.2 | official | |
| 38 | qwen2.5-max | 31.19 | official | |
| 39 | mistral-large-2 | 32.18 | official | |
| 39 | qwen2.5-72b-instruct | 32.18 | official | |
| 41 | chatgpt-4o | 38.12 | official | |
| 41 | mistral-medium-3 | 38.12 | official | |
| 43 | mistral-small-3 | 38.61 | official | |
| 44 | deepseek-v3 | 39.11 | official | |
| 45 | minimax-text-01 | 44.55 | official | |
| 46 | gemma-2-27b-it | 47.03 | official | |
| 47 | phi-4 | 52.48 | official | |
| 48 | nova-pro | 54.46 | official | |
| 49 | claude-3-haiku | 56.93 | official | |
| 50 | gpt-4o-mini | 60.89 | official | |
| 51 | claude-3.5-haiku | 65.84 | official | |
| 52 | gemma-3-27b-it | 66.34 | official |