Model profile
Vicuna 33B V1.3
Evidence summary
Vicuna 33B V1.3 has an estimated overall rank of #245; its 90% source-sensitivity interval is #192–#256. Its behavior-only rank is #250; company governance moves the combined estimate to #245. Published evidence spans 4 evals and 4 of 7 behavior components. Its strongest relative result is SALAD-Bench (mcq_representation_toxicity, #11 of 33); its weakest is SALAD-Bench (attack_enhanced_representation_toxicity, #33 of 33).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Hugging Face ↗lmsys/vicuna-33b-v1.3
- Official model page ↗Exact model document · Reviewed official Hugging Face owner · official repository
- Release source ↗global scoring family inherited
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Better | Source |
|---|---|---|---|---|---|
| Fake Alignment (FINE)multiple_choice_safe_decision_rate | #6 / 14 | 57.33 | ↑ higher | Source ↗official | |
| Fake Alignment (FINE)open_ended_safe_response_rate | #13 / 14 | 85.33 | ↑ higher | Source ↗official | |
| S-Evalbase_en_overall | #21 / 22 | 36.1 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_human_autonomy_integrity | #33 / 33 | 4.31 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_information_safety_harms | #29 / 33 | 4.23 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_malicious_use | #32 / 33 | 4 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_misinformation_harms | #30 / 33 | 4.61 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_representation_toxicity | #33 / 33 | 3.14 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_socioeconomic_harms | #33 / 33 | 3.9 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_human_autonomy_integrity | #31 / 33 | 58.12 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_information_safety_harms | #31 / 33 | 70.07 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_malicious_use | #31 / 33 | 46.84 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_misinformation_harms | #31 / 33 | 69.97 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_representation_toxicity | #31 / 33 | 52.78 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_socioeconomic_harms | #31 / 33 | 65.8 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_human_autonomy_integrity | #12 / 33 | 42.5 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_information_safety_harms | #15 / 33 | 36.94 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_malicious_use | #12 / 33 | 42.37 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_misinformation_harms | #15 / 33 | 39.76 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_representation_toxicity | #11 / 33 | 42.29 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_socioeconomic_harms | #13 / 33 | 40 | ↑ higher | Source ↗official | |
| SORRY-Benchavg | #43 / 51 | 0.56 | ↓ lower | Source ↗official |
