What Stanford actually built

On March 13 a group at Stanford's Center for Research on Foundation Models released Alpaca, a LLaMA 7B model fine-tuned to follow instructions. The training set is 52,000 instruction and response pairs. None of them were written by hand. The team started from 175 human-written seed pairs from the self-instruct paper, then prompted OpenAI's text-davinci-003 to generate more instructions and responses with those seeds as context.

The cost figures are the part everyone repeated. Generating the data cost less than $500 through the OpenAI API. Fine-tuning took about three hours on eight 80GB A100s, which the authors put at under $100 on most cloud providers. So the whole recipe, base model excluded, came in under $600.

The base model is the expensive ingredient and it was not theirs. LLaMA was released by Meta under a non-commercial research license, and the training data came out of a model whose terms forbid using it to build competitors. Alpaca is therefore non-commercial by construction, and the authors said so.

The evaluation behind the headline

The claim that got attention was that Alpaca behaves much like text-davinci-003. The evidence is a blind pairwise comparison run by the five student authors on the self-instruct evaluation set, which covers user-facing tasks like email drafting, social media posts, and productivity requests. Alpaca won 90 comparisons and text-davinci-003 won 89.

That is a real result and it is a narrow one. Five graders who built the model, one evaluation set drawn from the same distribution the training data was generated to cover, and a win count that is essentially a coin flip. The authors were upfront that the set is small and that the comparison says nothing about the many things it does not test.

They also listed failure modes in the release itself. Alpaca hallucinates, and the example they gave was the model confidently naming the wrong capital of Tanzania. It produces toxic output and stereotypes. It writes well-formed text that spreads misinformation. None of this was hidden, but none of it appeared in the screenshots that circulated.

The imitators arrive

Within a couple of weeks the recipe had been copied and enlarged. The one people know is Vicuna, released by LMSYS on March 30. It fine-tuned LLaMA 13B on around 70,000 conversations that users had shared publicly through ShareGPT, for a stated training cost of about $300. The pitch was a chatbot at 90 percent of ChatGPT quality.

That number came from asking GPT-4 to grade answers to 80 questions across eight categories, scoring for helpfulness, relevance, accuracy, and detail. Vicuna's total came to 92 percent of ChatGPT's under that grader. The Vicuna team wrote in the same post that this is not yet a rigorous or mature approach, since large models are prone to hallucinate, and that a standardised evaluation for chatbots remains an open problem. The caveat was accurate and it travelled far less well than the number.

What the demos measured

The most direct test of the imitation approach was published in May by Gudibande and colleagues at Berkeley. They fine-tuned models from 1.5 billion to 13 billion parameters on ChatGPT outputs, varying the amount of imitation data from 0.3 million to 150 million tokens, and had crowd workers compare the results against ChatGPT.

Crowd workers rated the imitation models as competitive with ChatGPT and much better than the base models at following instructions. On targeted evaluations of tasks outside the imitation data, the same models closed little or none of the gap between the base model and ChatGPT. The authors' phrasing was that imitation models pick up ChatGPT's style but not its factuality. Style is what a human sees in a short demo. Factuality is what a human notices after using the thing for a week.

A worked example makes this concrete. Ask an Alpaca-style model to write an email declining an invitation and the output is fluent and appropriately polite, because the training data is full of exactly that. Ask it a question that requires knowledge the 7B base model never learned well and it will answer in the same confident register, because the training data taught it the register and nothing else. The grader who judges tone will score both answers the same.

Where the cost actually went

So the $600 bought a large amount of surface behaviour and very little underlying capability. That is still a remarkable trade. Before March it was not obvious that a week of student effort could turn a raw 7B base model into something that holds a conversation. The cheapness is real and the models are useful for research on instruction following.

The expensive part turned out to be knowing what you built. A blind test by five authors on 179 comparisons cost almost nothing and told almost nothing. An 80-question GPT-4 grading run is cheap and, as its own authors said, unreliable. The Berkeley study, with human raters and targeted out-of-distribution tests, is the one that told us what was happening, and that is the kind of evaluation that does not get done in a weekend.

The Gudibande recommendation was to stop taking the shortcut and spend effort on better base models. We would add a second recommendation for anyone releasing a fine-tune this year. Publish the evaluation you ran, name the distribution it was drawn from, and say which of your claims it cannot support. Alpaca's authors did that. Most of the models that followed did not.

Sources

  1. Taori et al., Alpaca: A Strong, Replicable Instruction-Following Model (Stanford CRFM, March 13, 2023)
  2. Gudibande et al., The False Promise of Imitating Proprietary LLMs (arXiv 2305.15717)
  3. LMSYS, Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality (March 30, 2023)