Two answers to one problem

Language models are bad at arithmetic, out of date on facts and unaware of what day it is. The two papers that matter most on this at the moment agree on the fix, let the model call something outside itself, and disagree on almost everything else. Toolformer, from Timo Schick, Thomas Scialom and colleagues at Meta, teaches the model to use tools through fine-tuning on data it generates itself. ReAct, from Shunyu Yao and colleagues at Princeton and Google, teaches it through a prompt format and a loop, with no training at all.

How Toolformer gets its training data

The base model is GPT-J at 6.7B parameters and the corpus is a subset of CCNet. The trick is that the model annotates the corpus itself. For each tool the authors write a short prompt with a couple of examples of inserting an API call into text. The model then samples candidate call positions and candidate calls across the corpus, the calls are actually executed, and each call is kept only if prefixing the text with the call and its result reduces the model's loss on the following tokens by more than a threshold, compared to no call or to the call without its result.

That loss filter is the whole idea. It keeps calls that helped predict what came next and throws away the rest, with no human judging usefulness. At the loosest threshold the question answering tool survives in about 52,000 examples and the Wikipedia search in about 207,000, while the calculator is rare at 3,680 and machine translation at 3,156. The model is then fine-tuned on the corpus with those calls interleaved, which is otherwise the same text it saw before, so the authors argue it cannot lose its general language ability. Perplexity on WikiText and CCNet with calls disabled backs that up.

The five tools are a question answering system based on Atlas, a BM25 Wikipedia search, a four operation calculator, a calendar that returns the date, and NLLB for translation into English. The results are where the paper earns its attention. On the LAMA subsets Toolformer scores 33.8, 11.5 and 53.5 against 17.8, 4.9 and 31.9 for plain GPT-J, and above GPT-3 175B on all three. On the maths sets ASDiv, SVAMP and MAWPS it gets 40.4, 29.4 and 44.0 against 14.0, 10.0 and 19.8 for GPT-3, calling the calculator on 97.9 percent of examples. On open question answering it beats every GPT-J variant but stays well below GPT-3.

How ReAct runs a loop

ReAct does not train anything. It writes one or two examples into the prompt in which the model alternates between a thought, an action and an observation, and lets PaLM-540B continue the pattern. For question answering the actions are three operations on a Wikipedia interface, search for a topic, look up a string in the current page, and finish with an answer. The observation is whatever came back. The thought is free text where the model plans, notes what it has learned, and decides what to do next.

On HotpotQA the exact match numbers are 83.15 for ReAct against 78.60 for chain of thought alone and 74.40 for actions alone, and on FEVER 84.40 against 79.22 and 77.45. Combining ReAct with self-consistency chain of thought nudges those to 84.06 and 85.10. The larger gains are in the interactive environments. On ALFWorld, a text game about household tasks, ReAct reaches an 87.5 percent success rate against baselines around 71, and on WebShop 79.1 percent against about 58. The authors also fine-tune smaller PaLM models on ReAct traces and find they gain substantially, though they stay behind the 540B prompted model.

What the two papers assume

Toolformer assumes the tool set is small and each call is independent. The paper allows at most one API call per input at inference to stop the model looping on calls, and all training calls are sampled independently. The authors show the cost of this themselves on TempLAMA, where the best strategy would be to fetch the date and then query with it, and the model cannot, because chaining is prohibited and was never in the training data. The calendar tool is used on 0.2 percent of TempLAMA examples. Toolformer also cannot reformulate a query when the search returns junk, which the authors name as the reason it trails GPT-3 on open QA.

ReAct assumes retries are cheap and the environment is forgiving. The loop can search again when a result is bad, which is exactly what Toolformer lacks, but every step is a full pass through a 540B model, and the failure analysis groups the losses into reasoning errors, bad search results and hallucination. Nothing in the paper handles a tool that fails, times out or returns something the model cannot parse. The Wikipedia interface is deterministic and always answers.

Between them, the two papers fix the vocabulary that everyone building on this will use. A model emits a structured call, something executes it, the result comes back as text, and the model continues. Toolformer's contribution is the training signal for when to call. ReAct's contribution is the loop that lets a call's result change the next call. Neither paper has both, and the obvious next system does.

What we expect to break first

The few tools assumption goes first. Toolformer's scaling figure shows the ability to use tools only appearing around 775M parameters, and the loss filter produces very few examples for tools that are rarely useful in web text, 138 calculator calls at the strictest threshold. A system with dozens of tools will not get enough self-generated training data for most of them and will need a different way to learn when each applies. The cheap retries assumption goes second. ReAct on a 540B model with a loop of five or ten steps is expensive per question, and in any setting where the tool has side effects, a purchase or a file write, a retry is not free at all.

The experiment we would run is to fine-tune a GPT-J sized model on ReAct style traces with Toolformer style loss filtering applied to each step, so that the model learns both when a call helps and how to follow up on a result. The two papers make that a natural combination and neither set of authors has tried it.

Sources

  1. arXiv: Toolformer: Language Models Can Teach Themselves to Use Tools (Schick et al.)
  2. arXiv: ReAct: Synergizing Reasoning and Acting in Language Models (Yao et al.)