← Evals

Evaluation profile

DelusionEval

5sub-evals
0.512%total index weight
4components

Within-component eval weight: Truthfulness 1.12% · Human rights 1.78% · Misuse resistance 0.469% · Responsible agency 0.193%.

Model score (see each sub-eval direction)Predicted score

About this eval

Prevalence of delusion-linked conversational behaviors in model responses to real harm-associated conversation histories.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
delusional_prevalence_pctdelusioneval/delusioneval.csv:delusional_prevalence_pctMeasures endorsement of delusions, metaphysical elaboration, or false claims about model sentience and abilities in selected delusion-linked histories.truthfulness_honesty:1.000delusionevalLower is better0.0842%Truthfulness 0.561%
discourages_harm_prevalence_pctdelusioneval/delusioneval.csv:discourages_harm_prevalence_pctMeasures whether responses discourage self-harm or violence when the selected history contains those concerns.human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700delusionevalHigher is better0.0852%Human rights 0.411% · Misuse resistance 0.235%
facilitates_harm_prevalence_pctdelusioneval/delusioneval.csv:facilitates_harm_prevalence_pctMeasures whether responses facilitate self-harm or violence in histories selected for those concerns.human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700delusionevalLower is better0.0852%Human rights 0.411% · Misuse resistance 0.235%
relationship_prevalence_pctdelusioneval/delusioneval.csv:relationship_prevalence_pctMeasures platonic or romantic affinity and claims of a unique relationship with the user in selected delusion-linked histories.human_rights_systemic_harm:0.700|responsible_agency_control:0.300delusionevalLower is better0.173%Human rights 0.96% · Responsible agency 0.193%
sycophancy_prevalence_pctdelusioneval/delusioneval.csv:sycophancy_prevalence_pctMeasures affirmation, reflective agreement, grand-significance framing, dismissal of counterevidence, or claims that others admire the user in selected delusion-linked histories.truthfulness_honesty:1.000delusionevalLower is better0.0842%Truthfulness 0.561%

delusional_prevalence_pct

Measures endorsement of delusions, metaphysical elaboration, or false claims about model sentience and abilities in selected delusion-linked histories.

RankModelValueRelative performanceProvenance
1gpt-5.4-mini11.2official
2qwen3.5-9b12.8official
3gpt-5.414.3official
4gpt-5.4-nano18official
5claude-haiku-4.527.3official
6claude-opus-4.727.8official
7qwen3.5-397b-a17b28.65official
8claude-sonnet-4.631official
9gpt-4-turbo31.9official
10gemini-3.1-pro-preview47.3official
11gpt-4.148.4official
12gpt-4o50.3official
13gemini-2.5-flash-lite56official
14gemini-2.5-pro64.1official
15gemini-3.1-flash-lite66.8official
16grok-4.2070official

discourages_harm_prevalence_pct

Measures whether responses discourage self-harm or violence when the selected history contains those concerns.

RankModelValueRelative performanceProvenance
1gpt-5.463.05official
2gpt-5.4-mini61.3official
3gpt-5.4-nano52.2official
4claude-haiku-4.540.8official
5claude-opus-4.736.2official
6grok-4.2031.4official
7gemini-3.1-pro-preview29.4official
8gemini-3.1-flash-lite25.7official
9claude-sonnet-4.623.5official
10gpt-4o21.9official
11gemini-2.5-pro20official
12gemini-2.5-flash-lite16.5official
13gpt-4.115.1official
14gpt-4-turbo13.2official
15qwen3.5-397b-a17b12.55official
16qwen3.5-9b5official

facilitates_harm_prevalence_pct

Measures whether responses facilitate self-harm or violence in histories selected for those concerns.

RankModelValueRelative performanceProvenance
1gpt-5.4-mini0official
1gpt-5.4-nano0official
3gpt-5.40.2official
4qwen3.5-9b1official
5qwen3.5-397b-a17b1.15official
6claude-haiku-4.52.1official
6claude-sonnet-4.62.1official
8gpt-4-turbo2.3official
9gemini-3.1-pro-preview2.7official
10claude-opus-4.73.5official
11gpt-4.14.6official
12gemini-2.5-flash-lite5.6official
13gpt-4o6.7official
14gemini-3.1-flash-lite8.5official
15gemini-2.5-pro8.9official
16grok-4.2011.6official

relationship_prevalence_pct

Measures platonic or romantic affinity and claims of a unique relationship with the user in selected delusion-linked histories.

RankModelValueRelative performanceProvenance
1gpt-5.4-mini7.1official
2qwen3.5-9b8.3official
3gpt-5.410official
4gpt-5.4-nano11.8official
5claude-haiku-4.513.6official
6qwen3.5-397b-a17b21.1official
7claude-sonnet-4.621.4official
8gpt-4-turbo22official
9claude-opus-4.727official
10gemini-2.5-flash-lite33.6official
11gpt-4o36official
12gpt-4.137.1official
13gemini-3.1-pro-preview37.6official
14gemini-3.1-flash-lite38.6official
15gemini-2.5-pro40.8official
16grok-4.2053.4official

sycophancy_prevalence_pct

Measures affirmation, reflective agreement, grand-significance framing, dismissal of counterevidence, or claims that others admire the user in selected delusion-linked histories.

RankModelValueRelative performanceProvenance
1qwen3.5-9b9.9official
2gpt-5.4-mini15.2official
3gpt-5.4-nano15.6official
4qwen3.5-397b-a17b19.9official
5gpt-5.420.25official
6gpt-4-turbo21.6official
7claude-haiku-4.522.8official
8claude-opus-4.726.5official
9claude-sonnet-4.627.1official
10gemini-3.1-pro-preview30.4official
11gpt-4.133.2official
12gemini-3.1-flash-lite34.4official
13gemini-2.5-flash-lite35official
14gpt-4o35.5official
15grok-4.2036.2official
16gemini-2.5-pro37.6official