GAIA: questions humans get right 92 percent of the time and GPT-4 gets 15
GAIA turns the benchmark design problem around. Instead of finding questions that stump experts, it asks questions any careful person can answer and watches assistants fall over.
The inversion
Every hard benchmark of the last year has chased the same goal, questions that are difficult for humans. Graduate exams, competition maths, law bar questions. GAIA, released this week by a group at Meta, Hugging Face and AutoGPT, argues that this is the wrong direction. Tasks that are difficult for humans, the authors write, are not necessarily difficult for recent systems. Their proposal is the opposite. Questions that are conceptually simple for a person and require accurate execution of a long sequence of actions.
The headline numbers make the case. Human respondents score 92 percent overall. GPT-4 with plugins scores 15 percent. GPT-4 without tools does worse. Nothing about the questions is obscure. The example the paper leads with asks for the actual enrolment count of a clinical trial on H. pylori in acne patients listed on the NIH website for early 2018. The answer is 90. A person with a browser finds it in a few minutes. A language model has to browse, read a table, and not hallucinate a number.
Three levels and where the scores go
The 466 questions are split into three levels by how much a solver has to do. Level 1 needs no tools or at most one, in about five steps or fewer. Level 2 needs roughly five to ten steps combining several tools. Level 3 is the near-perfect general assistant case, arbitrarily long action sequences, multiple tools and access to the world. There are 146, 245 and 75 questions at each level.
Humans score 93.9, 91.8 and 87.3 percent across the three. GPT-4 with plugins scores 30.3, 9.7 and 0. Plain GPT-4 scores 9.1, 2.6 and 0. The shape of that table is the point. Human performance barely degrades with the number of steps, because people are good at chaining simple actions without losing the thread. Model performance collapses. Level 3 is a clean zero for both configurations.
The paper keeps 300 of the questions private and releases 166 with answers for a public leaderboard. That is the anti-contamination measure, and the authors are explicit that a benchmark whose answers can be found in plain text on the internet stops measuring anything once the next crawl runs.
How the questions were built
The design principles are stated up front. Real-world tasks that need reasoning, multimodality, browsing and tool use. Interpretability, in the sense that a non-expert can follow why an answer is right. Non-gameability, meaning verifiable answers that are not in pretraining data. And ease of use, with answers that are a number, a few words or a short list, scored by quasi exact match after normalisation. There is no model-as-judge, and the paper is pointed about why. Model-based evaluation relies on a more capable model than the one being tested, which is circular at the frontier.
Annotators wrote questions grounded in sources like Wikipedia, arXiv and GitHub, then two independent validators attempted each one. The instructions included making sure the answer does not exist on the internet in plain text and making sure it is unambiguous. Only 68 percent of questions passed validation on the first attempt. The rest were corrected or dropped. Each question cost about two hours of human time including validation, which is why there are 466 of them and not 10,000.
What GAIA measures that MMLU cannot
Multiple choice knowledge tests measure recall plus elimination. A model that has read the internet does well on them, and the gap between models at the top of MMLU is now within the benchmark's own error rate. GAIA measures something orthogonal. Can the system carry a plan across many steps, use tools without misreading their output, and commit to a single checkable answer at the end. Those are the properties that decide whether an assistant is useful, and they are the properties the current generation is worst at.
The 92 percent human figure is also a different kind of baseline from expert accuracy on a graduate exam. It says a careful non-expert gets nearly everything right. When a model reaches that number, the claim will be that it can do what a competent person with a browser can do, which is a claim with a meaning outside a leaderboard.
The weakness we see is narrowness of failure mode. A model that gets 15 percent on GAIA could be failing at browsing, at arithmetic on tables, at keeping state, or at formatting the final answer, and the score does not tell you which. The paper releases the questions with annotated steps, so the per-step analysis is possible, but nobody has done it yet.
What we expect
We expect GAIA to become the default eval for anything that calls itself an agent, because it is cheap to run, hard to game, and the human baseline is unarguable. We also expect Level 1 to be solved within a year by systems that combine a strong model with a decent browser and a retry loop. Level 3 is the one to watch. A zero is easy to improve on and hard to make meaningful.
The experiment we would run is to take the 166 public questions and log every tool call a system makes on its way to a wrong answer. Our guess is that most failures are one bad step early followed by confident execution of a wrong plan, and that the fix is in recovery rather than in raw capability. If that is right, GAIA scores will jump when someone builds a system that notices it is lost.
Sources
From the foundation