Model profile
LMSYSdeveloper
2023-06-22release date
#285 / 312overall rank
4eval lineages
Evidence summary
Vicuna 33B V1.3 has an estimated overall rank of #285; its 90% source-sensitivity interval is #220–#295. Its behavior-only rank is #292; company governance moves the combined estimate to #285. Published evidence spans 4 evals and 4 of 7 behavior components. Its strongest relative result is SALAD-Bench (mcq_representation_toxicity, #11 of 33); its weakest is SALAD-Bench (attack_enhanced_representation_toxicity, #33 of 33).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Hugging Face ↗lmsys/vicuna-33b-v1.3
- Official model page ↗Exact model document · Reviewed official Hugging Face owner · official repository
- Release source ↗global scoring family inherited
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Source |
|---|---|---|---|---|
| Fake Alignment (FINE)multiple_choice_safe_decision_rate | #6 / 14 | ↑57.33 | Source ↗official | |
| Fake Alignment (FINE)open_ended_safe_response_rate | #13 / 14 | ↑85.33 | Source ↗official | |
| S-Evalbase_en_overall | #21 / 22 | ↑36.1 | Source ↗official | |
| SALAD-Benchattack_enhanced_human_autonomy_integrity | #33 / 33 | ↑4.31 | Source ↗official | |
| SALAD-Benchattack_enhanced_information_safety_harms | #29 / 33 | ↑4.23 | Source ↗official | |
| SALAD-Benchattack_enhanced_malicious_use | #32 / 33 | ↑4 | Source ↗official | |
| SALAD-Benchattack_enhanced_misinformation_harms | #30 / 33 | ↑4.61 | Source ↗official | |
| SALAD-Benchattack_enhanced_representation_toxicity | #33 / 33 | ↑3.14 | Source ↗official | |
| SALAD-Benchattack_enhanced_socioeconomic_harms | #33 / 33 | ↑3.9 | Source ↗official | |
| SALAD-Benchbase_human_autonomy_integrity | #31 / 33 | ↑58.12 | Source ↗official | |
| SALAD-Benchbase_information_safety_harms | #31 / 33 | ↑70.07 | Source ↗official | |
| SALAD-Benchbase_malicious_use | #31 / 33 | ↑46.84 | Source ↗official | |
| SALAD-Benchbase_misinformation_harms | #31 / 33 | ↑69.97 | Source ↗official | |
| SALAD-Benchbase_representation_toxicity | #31 / 33 | ↑52.78 | Source ↗official | |
| SALAD-Benchbase_socioeconomic_harms | #31 / 33 | ↑65.8 | Source ↗official | |
| SALAD-Benchmcq_human_autonomy_integrity | #12 / 33 | ↑42.5 | Source ↗official | |
| SALAD-Benchmcq_information_safety_harms | #15 / 33 | ↑36.94 | Source ↗official | |
| SALAD-Benchmcq_malicious_use | #12 / 33 | ↑42.37 | Source ↗official | |
| SALAD-Benchmcq_misinformation_harms | #15 / 33 | ↑39.76 | Source ↗official | |
| SALAD-Benchmcq_representation_toxicity | #11 / 33 | ↑42.29 | Source ↗official | |
| SALAD-Benchmcq_socioeconomic_harms | #13 / 33 | ↑40 | Source ↗official | |
| SORRY-Benchavg | #43 / 51 | ↓0.56 | Source ↗official |
