What the paper did

Michael Bommarito and Daniel Martin Katz posted a short paper on December 29 in which they ran OpenAI's text-davinci-003 on a full practice Multistate Bar Examination bought from the National Conference of Bar Examiners. The MBE is the multiple-choice half of the bar exam, 200 questions across seven subjects plus 25 experimental items, and it is typically worth half of the overall score. They chose the NCBE's own practice exam, dated 2019 in the document, on the argument that a paid PDF was unlikely to be in the training data.

They tried seven prompt formats and a grid of sampling settings, 107 simulated exams in total, and found that one format mattered. Asking the model to rank its top three choices improved accuracy substantially over asking for a single answer, and they ran 41 exams with that prompt across parameter settings. They also tried fine-tuning on 200 unseen practice questions with and without explanations, six variants in all, and every fine-tuned model did worse than the base model, which they attribute to too little training data.

The numbers

The best configuration scored 50.3 percent overall against a 25 percent guessing baseline. By subject, evidence came in at 63 percent and torts at 62, which the authors count as passing rates for those subjects even though the NCBE-reported student averages are 65 and 71, while civil procedure scored 52, constitutional law 49, real property 45, contracts 45, and criminal law and procedure 35. The average student answers 68 percent correctly. The model trailed by about 17 points overall, with the gap negligible in evidence, torts and civil procedure and large in criminal law.

The result they highlight most is the ranking behaviour. When the model was asked for three choices, the correct answer was in its top two 71 percent of the time and in its top three 88 percent of the time. Since there are four options, an 88 percent top-three rate means the model could reliably identify the single wrong answer. The authors call this strong non-entailment performance, and it is the finding we would keep. The model was better at eliminating distractors than at choosing among the plausible remainder, which is a different skill from the one the exam claims to test.

A footnote is doing a lot of work. The MBE is one component, and a raw score in the high fifties to low sixties combined with adequate essays would be enough in a plurality of states. So 50.3 percent is not a pass, and the paper says so. What it says instead is that the results strongly suggest an LLM will pass the MBE component in the near future.

Why licensing exams looked ideal

It is easy to see why this framing spread. A professional exam is written by people who spend their careers on item design, it has a published human baseline broken down by category, it is refreshed regularly, and its pass mark carries a meaning outside machine learning. Compared to a benchmark assembled from crowdworkers, the MBE looks like a gift. The paper's own data section reads as an argument for exactly this, describing the years of study a human needs and the one in five who still fail.

The other appeal is narrative. Fifty percent on a held-out reading comprehension set is a number. Passing torts is a sentence anyone can repeat. The authors are careful about this. We are making a prediction about how the result will be used, and we expect every model release this year to come with an exam table.

What the exam does not measure

The first gap is the one the authors name. They have no insight into the model's head layers, they cannot say why the ranking prompt helped, and they cannot rule out that some legal material, including public sources, was in the training data even if this particular PDF was not. A held-out exam controls for memorising the answers. It does not control for having read the textbook.

The second gap is what a pass certifies. The bar exam is a screen for people who have already spent three years in law school and will then be supervised, insured and disciplined. Its questions are proxies for a competence that the surrounding institutions guarantee. When a model passes the proxy without the institutions, the pass tells you the proxy was learnable, which is not what the exam was designed to measure. The authors' entailment observation points the same way. Eliminating three wrong answers from a written vignette is a text skill, and the exam counts it as legal judgment because for humans the two have always come together.

What we would want from the follow-ups that are surely coming is the per-question analysis this paper could not do. Which items did the model miss, were they the ones that require applying a rule to an unusual fact pattern, and did the misses cluster in a way that a law professor would recognise as a specific gap? An exam score that comes with that breakdown is a measurement. One that comes without it is a press release with a decimal point.

Sources

  1. Bommarito and Katz, GPT Takes the Bar Exam (arXiv 2212.14402)