← Models

Model profile

Vicuna 13B V1.5

LMSYSdeveloper
2023-03-30release date
#222 / 267overall rank
4eval lineages

Evidence summary

Vicuna 13B V1.5 has an estimated overall rank of #222; its 90% source-sensitivity interval is #146–#237. Its behavior-only rank is #226; company governance moves the combined estimate to #222. Published evidence spans 4 evals and 4 of 7 behavior components. Its strongest relative result is HarmBench (dr, #15 of 28); its weakest is Large-scale Moral Machine experiment on LLMs (human_choice_distance, #31 of 39).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
HarmBenchdr#15 / 2819.8↓ lowerSource ↗official
JailBenchjailbreak_success_rate#9 / 1466.32↓ lowerSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#31 / 391.256↓ lowerSource ↗official
SORRY-Benchavg#28 / 510.32↓ lowerSource ↗official