← Evals

Evaluation profile

Shell

3sub-evals
0.186%Safety weight
1.33%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Susceptibility to implicit domain-specific risks in education, finance, and management.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
education_jsrshell/shell.csv:education_jsrMeasures whether the model resists jailbreak prompts framed as requests in the education domain.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000shellSafety: lower · Freedom: higher0.062%0.445%
finance_jsrshell/shell.csv:finance_jsrMeasures whether the model resists jailbreak prompts framed as requests in the finance domain.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000shellSafety: lower · Freedom: higher0.062%0.445%
management_jsrshell/shell.csv:management_jsrMeasures whether the model resists jailbreak prompts framed as requests in the management domain.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000shellSafety: lower · Freedom: higher0.062%0.445%

education_jsr

Measures whether the model resists jailbreak prompts framed as requests in the education domain.

RankModelValueRelative performanceProvenance
1claude-sonnet-40.28official
2gpt-50.364official
3gemini-2.5-pro0.418official
4kimi-k20.426official
5qwen3-235b-a22b0.492official
6doubao-seed-1.60.576official
7o30.608official
8llama-3.1-8b-instruct0.658official
9deepseek-r10.672official
10llama-4-maverick0.696official
11mistral-large-20.79official
12gpt-4o0.804official
12qwen2.5-7b-instruct0.804official
14grok-40.81official

finance_jsr

Measures whether the model resists jailbreak prompts framed as requests in the finance domain.

RankModelValueRelative performanceProvenance
1claude-sonnet-40.174official
2gpt-50.19official
3kimi-k20.22official
4qwen3-235b-a22b0.3official
5o30.328official
6gemini-2.5-pro0.4official
7grok-40.486official
8deepseek-r10.522official
9llama-3.1-8b-instruct0.6official
10doubao-seed-1.60.692official
11llama-4-maverick0.716official
12gpt-4o0.826official
13mistral-large-20.912official
14qwen2.5-7b-instruct0.914official

management_jsr

Measures whether the model resists jailbreak prompts framed as requests in the management domain.

RankModelValueRelative performanceProvenance
1claude-sonnet-40.17official
2kimi-k20.346official
3gpt-50.37official
4o30.482official
5gemini-2.5-pro0.502official
6qwen3-235b-a22b0.518official
7grok-40.596official
8doubao-seed-1.60.616official
9deepseek-r10.682official
10llama-3.1-8b-instruct0.724official
11llama-4-maverick0.844official
12gpt-4o0.872official
13mistral-large-20.92official
14qwen2.5-7b-instruct0.938official