Evaluation profile
FinEval Financial Security Knowledge
1sub-evals
0.0609%total index weight
1components
Within-component eval weight: Misuse resistance 0.609%.
Model score (higher is better)Predicted score
About this eval
Knowledge of financial cybersecurity and vulnerability topics.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| financial_security_accuracy_pctfineval/fineval-financial-security-knowledge.csv:financial_security_accuracy_pctMeasures cybersecurity knowledge in financial contexts. | ordinary_harm_misuse_resistance:1.000fineval-safety | Higher is better | 0.0609% | Misuse resistance 0.609% |
financial_security_accuracy_pct
Measures cybersecurity knowledge in financial contexts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o | 81.8 | official | |
| 1 | qwen2.5-72b-instruct | 81.8 | official | |
| 3 | gpt-4o-mini | 79.1 | official | |
| 4 | claude-3.5-sonnet | 78.1 | official | |
| 5 | gemini-1.5-pro | 77.8 | official | |
| 6 | gemini-1.5-flash | 77.5 | official | |
| 7 | yi-1.5-34b-chat | 76 | official | |
| 8 | xuanyuan3-70b-chat | 74.4 | official | |
| 9 | internlm2.5-20b-chat | 74.1 | official | |
| 10 | glm-4-9b-chat | 73.1 | official | |
| 10 | internlm2-chat-20b | 73.1 | official | |
| 12 | qwen-2.5-7b-instruct | 71.7 | official | |
| 13 | yi-1.5-9b-chat | 71.4 | official | |
| 14 | xuanyuan2-70b-chat | 68 | official | |
| 15 | cfgpt2-7b | 65.1 | official | |
| 16 | baichuan-2-13b-chat | 61.6 | official | |
| 17 | chatglm3-6b | 48.2 | official | |
| 18 | disc-finllm | 25.2 | official | |
| 19 | fingpt-v3.1 | 22.7 | official |