What was released

Yesterday a team from UC Berkeley, CMU, Stanford, UC San Diego, and MBZUAI released Vicuna-13B, a LLaMA fine-tune trained on roughly 70,000 user-shared conversations collected from ShareGPT. The training cost about 300 dollars, ran on eight A100s in about a day using SkyPilot managed spot instances, extended the context to 2,048 tokens, and adjusted the loss so that multi-turn conversations were weighted properly.

The headline of the post is that Vicuna-13B achieves more than 90 percent of the quality of ChatGPT and Bard. The asterisk on that number is the part we care about. The quality figure is not a benchmark score. It comes from asking GPT-4 to grade the answers.

How the judging worked

The team wrote 80 questions across eight categories, including Fermi problems, roleplay scenarios, and coding and math tasks. They collected answers from five models, LLaMA, Alpaca, ChatGPT, Bard, and Vicuna, and asked GPT-4 to rate each answer on helpfulness, relevance, accuracy, and detail. Summed up, Vicuna's total score came to 92 percent of ChatGPT's. In pairwise terms, GPT-4 preferred Vicuna over Alpaca and LLaMA in more than 90 percent of cases, and rated Vicuna's answer as better than or equal to ChatGPT's on 45 percent of the questions.

Two different numbers are doing work here and they are easy to conflate. The 92 percent is a ratio of summed scores, so a model that is a little worse on every question would score near 92 too. The 45 percent is a win-or-tie rate. Both come from the same 80 questions and the same judge.

The footnote that everyone will forget

The post says, in plain words, that the framework is not yet a rigorous or mature approach, because large language models are prone to hallucinate. It also notes that the judge and the models being judged share weaknesses, with basic math and limited coding ability named for both. That is an honest caveat and it is placed where people will skip past it.

The reasons the caveat matters are concrete. The judge is a model in the same family as one of the contestants, so any stylistic preference GPT-4 has for ChatGPT-shaped answers goes straight into the score. Eighty questions is a small sample and no uncertainty is reported. The categories were chosen by the same people who built the model. And a judge that cannot do the math itself is grading whether the math is right.

Why this method will spread anyway

None of that will stop the method spreading, and we think the reasons are worth stating so we know what we are signing up for. Human evaluation of open-ended chat answers is slow and costly. Existing benchmarks were built for classification and short answers and do not measure whether a chatbot is pleasant and useful. GPT-4 is available through an API, gives a number in seconds, and its judgements look sensible when you spot check them.

The open-model community also has a specific need right now. LLaMA weights leaked in early March, Alpaca showed that a cheap instruction tune changes behaviour a lot, and a stream of fine-tunes is coming. Everyone wants to know whether theirs is better than the last one. A judge that costs a few dollars per comparison answers that question at the speed the field is moving. Rigour will have to catch up later. The risk is that a number invented as a rough guide gets quoted as a measurement, and that a leaderboard grows up around a judge nobody has calibrated.

What we would want before trusting a judge score

If GPT-4-as-judge becomes the default, the minimum we would ask of anyone reporting a score is a few things the Vicuna team did not have time for. Report agreement between the judge and human raters on a subset. Swap the order of the two answers in pairwise comparisons and check whether the preference flips. Hold the question set fixed and public so the numbers are comparable across releases. Report a confidence interval, even a crude bootstrap over the 80 questions.

The Vicuna post already acknowledges that the model is weak at reasoning and mathematics, sometimes fails to identify itself accurately, and has not been tuned for safety. Those are the honest limits of a 300 dollar fine-tune. The evaluation is the part that will be copied, and we would rather the caveat travelled with it.

Sources

  1. LMSYS, Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality