Evaluation profile
Shell
3sub-evals
0.239%total index weight
1components
Within-component eval weight: Misuse resistance 2.39%.
Model score (lower is better)Predicted score
About this eval
Susceptibility to implicit domain-specific risks in education, finance, and management.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| education_jsrshell/shell.csv:education_jsrMeasures whether the model resists jailbreak prompts framed as requests in the education domain. | ordinary_harm_misuse_resistance:1.000shell | Lower is better | 0.0798% | Misuse resistance 0.798% |
| finance_jsrshell/shell.csv:finance_jsrMeasures whether the model resists jailbreak prompts framed as requests in the finance domain. | ordinary_harm_misuse_resistance:1.000shell | Lower is better | 0.0798% | Misuse resistance 0.798% |
| management_jsrshell/shell.csv:management_jsrMeasures whether the model resists jailbreak prompts framed as requests in the management domain. | ordinary_harm_misuse_resistance:1.000shell | Lower is better | 0.0798% | Misuse resistance 0.798% |
education_jsr
Measures whether the model resists jailbreak prompts framed as requests in the education domain.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4 | 0.28 | official | |
| 2 | gpt-5 | 0.364 | official | |
| 3 | gemini-2.5-pro | 0.418 | official | |
| 4 | kimi-k2 | 0.426 | official | |
| 5 | qwen3-235b-a22b | 0.492 | official | |
| 6 | doubao-seed-1.6 | 0.576 | official | |
| 7 | o3 | 0.608 | official | |
| 8 | llama-3.1-8b-instruct | 0.658 | official | |
| 9 | deepseek-r1 | 0.672 | official | |
| 10 | llama-4-maverick | 0.696 | official | |
| 11 | mistral-large-2 | 0.79 | official | |
| 12 | gpt-4o | 0.804 | official | |
| 12 | qwen-2.5-7b-instruct | 0.804 | official | |
| 14 | grok-4 | 0.81 | official |
finance_jsr
Measures whether the model resists jailbreak prompts framed as requests in the finance domain.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4 | 0.174 | official | |
| 2 | gpt-5 | 0.19 | official | |
| 3 | kimi-k2 | 0.22 | official | |
| 4 | qwen3-235b-a22b | 0.3 | official | |
| 5 | o3 | 0.328 | official | |
| 6 | gemini-2.5-pro | 0.4 | official | |
| 7 | grok-4 | 0.486 | official | |
| 8 | deepseek-r1 | 0.522 | official | |
| 9 | llama-3.1-8b-instruct | 0.6 | official | |
| 10 | doubao-seed-1.6 | 0.692 | official | |
| 11 | llama-4-maverick | 0.716 | official | |
| 12 | gpt-4o | 0.826 | official | |
| 13 | mistral-large-2 | 0.912 | official | |
| 14 | qwen-2.5-7b-instruct | 0.914 | official |
management_jsr
Measures whether the model resists jailbreak prompts framed as requests in the management domain.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-sonnet-4 | 0.17 | official | |
| 2 | kimi-k2 | 0.346 | official | |
| 3 | gpt-5 | 0.37 | official | |
| 4 | o3 | 0.482 | official | |
| 5 | gemini-2.5-pro | 0.502 | official | |
| 6 | qwen3-235b-a22b | 0.518 | official | |
| 7 | grok-4 | 0.596 | official | |
| 8 | doubao-seed-1.6 | 0.616 | official | |
| 9 | deepseek-r1 | 0.682 | official | |
| 10 | llama-3.1-8b-instruct | 0.724 | official | |
| 11 | llama-4-maverick | 0.844 | official | |
| 12 | gpt-4o | 0.872 | official | |
| 13 | mistral-large-2 | 0.92 | official | |
| 14 | qwen-2.5-7b-instruct | 0.938 | official |