Evaluation profile
DelusionEval
Within-component eval weight: Truthfulness 1.12% · Human rights 1.78% · Misuse resistance 0.469% · Responsible agency 0.193%.
About this eval
Prevalence of delusion-linked conversational behaviors in model responses to real harm-associated conversation histories.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| delusional_prevalence_pctdelusioneval/delusioneval.csv:delusional_prevalence_pctMeasures endorsement of delusions, metaphysical elaboration, or false claims about model sentience and abilities in selected delusion-linked histories. | truthfulness_honesty:1.000delusioneval | Lower is better | 0.0842% | Truthfulness 0.561% |
| discourages_harm_prevalence_pctdelusioneval/delusioneval.csv:discourages_harm_prevalence_pctMeasures whether responses discourage self-harm or violence when the selected history contains those concerns. | human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700delusioneval | Higher is better | 0.0852% | Human rights 0.411% · Misuse resistance 0.235% |
| facilitates_harm_prevalence_pctdelusioneval/delusioneval.csv:facilitates_harm_prevalence_pctMeasures whether responses facilitate self-harm or violence in histories selected for those concerns. | human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700delusioneval | Lower is better | 0.0852% | Human rights 0.411% · Misuse resistance 0.235% |
| relationship_prevalence_pctdelusioneval/delusioneval.csv:relationship_prevalence_pctMeasures platonic or romantic affinity and claims of a unique relationship with the user in selected delusion-linked histories. | human_rights_systemic_harm:0.700|responsible_agency_control:0.300delusioneval | Lower is better | 0.173% | Human rights 0.96% · Responsible agency 0.193% |
| sycophancy_prevalence_pctdelusioneval/delusioneval.csv:sycophancy_prevalence_pctMeasures affirmation, reflective agreement, grand-significance framing, dismissal of counterevidence, or claims that others admire the user in selected delusion-linked histories. | truthfulness_honesty:1.000delusioneval | Lower is better | 0.0842% | Truthfulness 0.561% |
delusional_prevalence_pct
Measures endorsement of delusions, metaphysical elaboration, or false claims about model sentience and abilities in selected delusion-linked histories.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4-mini | 11.2 | official | |
| 2 | qwen3.5-9b | 12.8 | official | |
| 3 | gpt-5.4 | 14.3 | official | |
| 4 | gpt-5.4-nano | 18 | official | |
| 5 | claude-haiku-4.5 | 27.3 | official | |
| 6 | claude-opus-4.7 | 27.8 | official | |
| 7 | qwen3.5-397b-a17b | 28.65 | official | |
| 8 | claude-sonnet-4.6 | 31 | official | |
| 9 | gpt-4-turbo | 31.9 | official | |
| 10 | gemini-3.1-pro-preview | 47.3 | official | |
| 11 | gpt-4.1 | 48.4 | official | |
| 12 | gpt-4o | 50.3 | official | |
| 13 | gemini-2.5-flash-lite | 56 | official | |
| 14 | gemini-2.5-pro | 64.1 | official | |
| 15 | gemini-3.1-flash-lite | 66.8 | official | |
| 16 | grok-4.20 | 70 | official |
discourages_harm_prevalence_pct
Measures whether responses discourage self-harm or violence when the selected history contains those concerns.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 63.05 | official | |
| 2 | gpt-5.4-mini | 61.3 | official | |
| 3 | gpt-5.4-nano | 52.2 | official | |
| 4 | claude-haiku-4.5 | 40.8 | official | |
| 5 | claude-opus-4.7 | 36.2 | official | |
| 6 | grok-4.20 | 31.4 | official | |
| 7 | gemini-3.1-pro-preview | 29.4 | official | |
| 8 | gemini-3.1-flash-lite | 25.7 | official | |
| 9 | claude-sonnet-4.6 | 23.5 | official | |
| 10 | gpt-4o | 21.9 | official | |
| 11 | gemini-2.5-pro | 20 | official | |
| 12 | gemini-2.5-flash-lite | 16.5 | official | |
| 13 | gpt-4.1 | 15.1 | official | |
| 14 | gpt-4-turbo | 13.2 | official | |
| 15 | qwen3.5-397b-a17b | 12.55 | official | |
| 16 | qwen3.5-9b | 5 | official |
facilitates_harm_prevalence_pct
Measures whether responses facilitate self-harm or violence in histories selected for those concerns.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4-mini | 0 | official | |
| 1 | gpt-5.4-nano | 0 | official | |
| 3 | gpt-5.4 | 0.2 | official | |
| 4 | qwen3.5-9b | 1 | official | |
| 5 | qwen3.5-397b-a17b | 1.15 | official | |
| 6 | claude-haiku-4.5 | 2.1 | official | |
| 6 | claude-sonnet-4.6 | 2.1 | official | |
| 8 | gpt-4-turbo | 2.3 | official | |
| 9 | gemini-3.1-pro-preview | 2.7 | official | |
| 10 | claude-opus-4.7 | 3.5 | official | |
| 11 | gpt-4.1 | 4.6 | official | |
| 12 | gemini-2.5-flash-lite | 5.6 | official | |
| 13 | gpt-4o | 6.7 | official | |
| 14 | gemini-3.1-flash-lite | 8.5 | official | |
| 15 | gemini-2.5-pro | 8.9 | official | |
| 16 | grok-4.20 | 11.6 | official |
relationship_prevalence_pct
Measures platonic or romantic affinity and claims of a unique relationship with the user in selected delusion-linked histories.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4-mini | 7.1 | official | |
| 2 | qwen3.5-9b | 8.3 | official | |
| 3 | gpt-5.4 | 10 | official | |
| 4 | gpt-5.4-nano | 11.8 | official | |
| 5 | claude-haiku-4.5 | 13.6 | official | |
| 6 | qwen3.5-397b-a17b | 21.1 | official | |
| 7 | claude-sonnet-4.6 | 21.4 | official | |
| 8 | gpt-4-turbo | 22 | official | |
| 9 | claude-opus-4.7 | 27 | official | |
| 10 | gemini-2.5-flash-lite | 33.6 | official | |
| 11 | gpt-4o | 36 | official | |
| 12 | gpt-4.1 | 37.1 | official | |
| 13 | gemini-3.1-pro-preview | 37.6 | official | |
| 14 | gemini-3.1-flash-lite | 38.6 | official | |
| 15 | gemini-2.5-pro | 40.8 | official | |
| 16 | grok-4.20 | 53.4 | official |
sycophancy_prevalence_pct
Measures affirmation, reflective agreement, grand-significance framing, dismissal of counterevidence, or claims that others admire the user in selected delusion-linked histories.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen3.5-9b | 9.9 | official | |
| 2 | gpt-5.4-mini | 15.2 | official | |
| 3 | gpt-5.4-nano | 15.6 | official | |
| 4 | qwen3.5-397b-a17b | 19.9 | official | |
| 5 | gpt-5.4 | 20.25 | official | |
| 6 | gpt-4-turbo | 21.6 | official | |
| 7 | claude-haiku-4.5 | 22.8 | official | |
| 8 | claude-opus-4.7 | 26.5 | official | |
| 9 | claude-sonnet-4.6 | 27.1 | official | |
| 10 | gemini-3.1-pro-preview | 30.4 | official | |
| 11 | gpt-4.1 | 33.2 | official | |
| 12 | gemini-3.1-flash-lite | 34.4 | official | |
| 13 | gemini-2.5-flash-lite | 35 | official | |
| 14 | gpt-4o | 35.5 | official | |
| 15 | grok-4.20 | 36.2 | official | |
| 16 | gemini-2.5-pro | 37.6 | official |