← Evals

Evaluation profile

ThaiSafetyBench

1sub-evals
0.0667%total index weight
1components

Within-component eval weight: Misuse resistance 0.667%.

Model score (higher is better)Predicted score

About this eval

Resistance to Thai-language malicious prompts, including culturally contextualized attacks.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
safety_scorethaisafetybench/thaisafetybench.csv:safety_scoreMeasures resistance to harmful Thai-language prompt attacks.ordinary_harm_misuse_resistance:1.000do-not-answer/do-not-answer.csvHigher is better0.0667%Misuse resistance 0.667%

safety_score

Measures resistance to harmful Thai-language prompt attacks.

RankModelValueRelative performanceProvenance
1gpt-595.57official
2claude-sonnet-4.590.25official
3seallms-v3-7b-chat90.17official
4qwen2.5-72b-instruct89.01official
5openthaigpt1.5-72b-instruct87.66official
6llama-sea-lion-v3-70b-it87.3official
7qwen-2.5-7b-instruct85.57official
8seallms-v3-1.5b-chat85.39official
9gpt-4o83.96official
10openthaigpt1.5-7b-instruct83.91official
11llama-3.3-70b-instruct83.13official
12llama-sea-lion-v3-8b-it83.1official
13gemma-3-12b79.6official
14llama-3.1-70b-instruct75.51official
15llama-3.2-3b73.92official
16gemma-3-4b71.89official
17llama-3.1-8b-instruct71.76official
18llama-3.2-1b-instruct62.34official