What was released

On April 12 Databricks released Dolly 2.0, a 12 billion parameter model fine-tuned from EleutherAI's Pythia-12b, and a dataset called databricks-dolly-15k. The model is licensed for commercial use. The dataset is licensed under Creative Commons Attribution-ShareAlike 3.0. The company says the fine-tuning cost less than 30 dollars, and the original Dolly a few weeks earlier had already shown that a small budget on an old base model produces something that follows instructions.

The model is not the story. The dataset card lists 15,011 records, 13.1 megabytes, American English, written by Databricks employees between March and April. That is the first instruction-following dataset we know of that a company can train on and ship a product with, without a lawyer asking where the examples came from.

Why the licence is the whole point

Every open instruction-tuned model released since March has the same problem. Alpaca, Koala, GPT4All and Vicuna were all fine-tuned on outputs from ChatGPT or from a model trained on ChatGPT outputs, and OpenAI's terms prohibit using those outputs to build competing models. The Databricks post names all four. Whatever you think of the enforceability of that clause, no company wants to be the test case. So the open models were open for research and closed for everything else.

Databricks took the obvious route around this, which was also the expensive one. Instead of asking a model to write the examples, they asked over 5,000 employees. The post says it was run as a contest with a nightly leaderboard and prizes for the top 20 contributors. The inspiration was OpenAI's InstructGPT paper, which trained on roughly 13,000 human demonstrations, so the target size was set by what had been shown to work rather than by what was easy.

What the annotators were told

The dataset card is more revealing than the blog post. Annotators got short guidelines with examples for each category and two hard rules. They were not to use information from any source on the web except Wikipedia. And they were not to use generative AI to write either the instructions or the responses. The second rule is the one that makes the dataset legally clean, and it is also the one that makes it slow, because it removes the shortcut every other group had taken.

The categories are open question answering, closed question answering over a supplied Wikipedia passage, information extraction from a reference text, summarisation of a Wikipedia paragraph, classification, brainstorming, creative writing, and a free-form bucket. The blog post counts seven and the card counts eight. Either way, the Wikipedia-grounded tasks are a large share, which means the dataset is strong on reading a passage and weak on anything that needs knowledge the annotator would have had to look up elsewhere.

The hidden cost of not distilling

The Alpaca approach costs a few hundred dollars in API calls and an afternoon. The Dolly approach cost the attention of thousands of employees for weeks, plus whatever the prizes were, plus the organisational effort of running a contest. Databricks can absorb that because the dataset is a marketing asset for a company that sells data infrastructure. A university lab could not have done it, and neither could most startups.

There is also a quality cost that cuts the other way. The Databricks post argues that human-written examples avoid the hallucinations that distilled datasets inherit from the teacher model. That is probably true for the closed question answering tasks where the answer is in the supplied passage. The card is honest that annotators may not all be native English speakers and that their demographics reflect the company's workforce, which is a way of saying the dataset knows a lot about the concerns of software engineers at a data company.

What distillation had been hiding is that the cheap route to an instruction dataset runs through a closed model's licence, and that the alternative requires either money or a community. The Dolly dataset is a proof that the alternative works at the 15,000 example scale. It says nothing yet about whether it works at the scale the closed labs are operating at.

What we would do with it

The most useful experiment is the one Databricks did not run. Take the same base model, fine-tune it once on dolly-15k and once on a distilled dataset of the same size and category mix, and evaluate both with human raters on held-out prompts. The blog post says Dolly 2.0 is comparable to Alpaca despite the smaller dataset, but that comparison changes base model and dataset at the same time.

The second thing we would want is for someone to keep going. Fifteen thousand examples from one company is a start. If a few more organisations ran the same contest under the same licence and pooled the results, the open community would have a human-written instruction set that no terms of service can touch. The dataset licence requires share-alike, so anything built on it stays open. That was a deliberate choice and it is the part of this release we expect to matter in a year.

Sources

  1. Databricks, Free Dolly: Introducing the World's First Truly Open Instruction-Tuned LLM (April 12, 2023)
  2. databricks-dolly-15k dataset card on Hugging Face