← Evals

Evaluation profile

Vigil Mental Health Safety

1sub-evals
0.566%Safety weight
1.28%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Mental-health conversational safety when assisting vulnerable users across simulated multi-turn scenarios.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
overall_scorevigil-eval/vigil-eval.csv:overall_scoreMeasures whether an assistant maintains safe boundaries rather than colluding with or escalating vulnerable users during extended mental-health conversations.Safety: human_rights_systemic_harm:0.450|ordinary_harm_misuse_resistance:0.550 · Freedom: user_direction_following:1.000vigil-evalSafety: higher · Freedom: lower0.566%1.28%

overall_score

Measures whether an assistant maintains safe boundaries rather than colluding with or escalating vulnerable users during extended mental-health conversations.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.683official
2gpt-5.478official
3claude-sonnet-4.575official
4claude-haiku-4.574official
5gpt-5.272official
6gpt-5.3-chat71official
7gpt-5.158official
8claude-sonnet-453official
9gpt-551official
10gemini-3.1-pro-preview48official
11gemini-3-flash-preview47official
12gemma-4-31b-it46official
13gemini-3.1-flash-lite38official
13gemma-4-26b-a4b-it38official
15deepseek-v3.235official
15gpt-oss-120b35official
17gpt-oss-20b31official
18grok-4.1-fast30official
19gemini-2.5-flash28official
19llama-4-maverick28official
21gpt-4o27official
22gpt-4.126official
23llama-4-scout20official