Evaluation profile
Confabulations
1sub-evals
0.735%Safety weight
0%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Propensity to confabulate answers rather than identify insufficient knowledge or false premises.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| confabulation_rateconfabulations/confabulations.csv:confabulation_rateMeasures how often the model invents an answer when the prompt has a false premise or cannot be answered from reliable knowledge. | Safety: truthfulness_honesty:1.000confabulations | Safety: lower | 0.735% | — |
confabulation_rate
Measures how often the model invents an answer when the prompt has a false premise or cannot be answered from reliable knowledge.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4 | 2.723 | official | |
| 2 | claude-opus-4.1 | 3.218 | official | |
| 3 | claude-sonnet-4 | 3.96 | official | |
| 3 | gemini-2.5-pro-exp | 3.96 | official | |
| 3 | grok-4 | 3.96 | official | |
| 6 | gemini-2.5-flash | 4.455 | official | |
| 7 | gemini-2.5-pro | 4.95 | official | |
| 8 | glm-4.5 | 7.921 | official | |
| 9 | grok-3-mini | 8.911 | official | |
| 10 | o1 | 10.89 | official | |
| 11 | gpt-4.5-preview | 11.88 | official | |
| 12 | claude-3.5-sonnet | 12.87 | official | |
| 12 | qwen3-30b-a3b | 12.87 | official | |
| 14 | llama-3.1-405b-instruct | 14.36 | official | |
| 15 | deepseek-r1 | 15.1 | official | |
| 16 | gemini-2.0-pro-exp | 15.84 | official | |
| 17 | claude-3.7-sonnet | 16.58 | official | |
| 18 | gemini-1.5-pro | 16.83 | official | |
| 19 | grok-3 | 17.82 | official | |
| 19 | llama-3.3-70b-instruct | 17.82 | official | |
| 21 | o1-preview | 18.32 | official | |
| 22 | qwen3-235b-a22b | 18.56 | official | |
| 23 | gemini-2.0-flash | 19.55 | official | |
| 24 | gpt-oss-120b | 23.27 | official | |
| 25 | o3-pro | 23.38 | official | |
| 26 | gpt-4o | 24.26 | official | |
| 27 | o3 | 24.75 | official | |
| 28 | ernie-4.5-300b-a47b | 25.25 | official | |
| 28 | qwq-32b | 25.25 | official | |
| 30 | grok-2 | 25.74 | official | |
| 31 | o1-mini | 26.24 | official | |
| 32 | gpt-4-turbo | 26.73 | official | |
| 32 | o4-mini | 26.73 | official | |
| 34 | claude-3-opus | 28.22 | official | |
| 34 | llama-4-maverick | 28.22 | official | |
| 36 | o3-mini | 28.96 | official | |
| 37 | kimi-k2 | 30.2 | official | |
| 38 | qwen2.5-max | 31.19 | official | |
| 39 | mistral-large-2 | 32.18 | official | |
| 39 | qwen2.5-72b-instruct | 32.18 | official | |
| 41 | chatgpt-4o | 38.12 | official | |
| 41 | mistral-medium-3 | 38.12 | official | |
| 43 | mistral-small-3 | 38.61 | official | |
| 44 | deepseek-v3 | 39.11 | official | |
| 45 | minimax-text-01 | 44.55 | official | |
| 46 | gemma-2-27b-it | 47.03 | official | |
| 47 | phi-4 | 52.48 | official | |
| 48 | nova-pro | 54.46 | official | |
| 49 | claude-3-haiku | 56.93 | official | |
| 50 | gpt-4o-mini | 60.89 | official | |
| 51 | claude-3.5-haiku | 65.84 | official | |
| 52 | gemma-3-27b-it | 66.34 | official |