← Evals

Evaluation profile

AIMS Safety-Classifier Competence

1sub-evals
0.0926%total index weight
1components

Within-component eval weight: Misuse resistance 0.926%.

Model score (higher is better)Predicted score

About this eval

Harmful-prompt safety-classification competence.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
average_harmful_f1aims/aims.csv:average_harmful_f1Measures harmful-prompt classification competence.ordinary_harm_misuse_resistance:1.000aims-safety-classifier-competenceHigher is better0.0926%Misuse resistance 0.926%

average_harmful_f1

Measures harmful-prompt classification competence.

RankModelValueRelative performanceProvenance
1gpt-5.40.815official
2nemotron-content-safety-reasoning-4b0.809official
3gpt-oss-safeguard-120b0.807official
4guardreasoner-8b0.804official
4wildguard-7b0.804official
6gpt-oss-120b0.803official
7gemma-3-12b0.802official
8claude-sonnet-4.60.782official
9llama-3.1-8b-instruct0.749official
10shieldgemma-27b0.709official
11llama-guard-4-12b0.691official