The top 10 percent on the bar exam: reading the GPT-4 technical report's exam table
The GPT-4 report's abstract says the model passed a simulated bar exam with a score around the top 10 percent of test takers. Percentiles are only meaningful against a named population, and the population behind that figure is the one most likely to flatter the model. Notes on how to read the exam table, with an addendum on the re-analysis that put the number closer to the 48th percentile of people who passed.
The sentence everyone quoted
The GPT-4 technical report is a document that declines to say how the model was built, so the exam results have become the thing people repeat. The abstract's example is that GPT-4 passes a simulated bar exam with a score around the top 10 percent of test takers, and that it achieves human-level performance on various professional and academic benchmarks. Within a week that had been shortened in most coverage to GPT-4 scoring in the 90th percentile on the bar.
We want to take the claim seriously rather than dismiss it, because passing the Uniform Bar Examination with a real scaled score is a genuine result and the report's authors did not invent it. What we want to question is the percentile, since a percentile is a statement about a population, and the report does not say which population.
What a percentile requires
A score at the 90th percentile means 90 percent of some group scored lower. Change the group and the percentile moves without the score moving at all. For a licensing exam there are several natural groups: everyone who sat a given administration, first-time takers, repeat takers, and the subset who passed. These differ a lot. Repeat takers are, almost by definition, people who scored below the pass line last time. A percentile computed against a sitting heavy with repeaters will be higher than the same score's percentile against first-timers, and much higher than against people who passed.
The report gives a percentile without a denominator, and that is the whole problem. If you are comparing a model to a human professional, the meaningful reference group is people who do the job, which for lawyers means people who passed. If you are comparing it to a student sitting the exam for the first time, use first-time takers. Comparing it against a sitting dominated by people retaking the exam answers a question nobody was asking.
What a working researcher should do with the table
Three habits. First, when a paper reports a percentile, look for the population and the administration date, and if neither is given, treat the number as a rank against an unknown group and do not carry it into your own writing. Second, separate the components. A bar exam has a multiple-choice section and essay sections, and a model that reads well on the multiple choice may do very differently when a human grades free text against a rubric. The report gives one combined percentile, which hides that split.
Third, ask who graded the essays and how. Multiple-choice scoring is mechanical. Essay grading on a licensing exam is done by trained graders against the exam board's rubric, and a simulated grading by someone else is a different measurement. The model still passed. The number that travelled was the least interpretable one in the report.
Addendum, 2024: the re-analysis
A paper in Artificial Intelligence and Law has now done the work the report left undone, and it is worth appending the numbers here because they change the picture. The author replicates GPT-4's multiple-choice scaled score of 158 on the MBE, so the underlying result stands. The 90th percentile figure, though, appears to trace to a February Illinois administration, and about 70 percent of February takers are repeaters who failed the previous July and who score substantially lower than first-timers. That is exactly the flattering denominator we were worried about.
Against first-time takers, the paper estimates GPT-4's overall UBE score at roughly the 62nd percentile. Against people who passed, roughly the 48th. The split by component is the most useful part. On the MBE multiple choice the model sits around the 79th percentile of first-timers and the 69th of passers. On the essay and performance-test sections it is around the 42nd percentile of first-timers and the 15th of passers. So the model is a strong multiple-choice taker and a below-median essay writer relative to newly licensed lawyers, and the single headline number averaged those into something that described neither.
What we would like to see next is for any exam-based claim to come with the reference population, the administration, and per-component percentiles as a matter of course. The bar exam result was real. The 90th percentile was a choice of denominator.
Sources
From the foundation