A benchmark that was waiting

PlanBench has been around since 2022, and for two years its main finding was monotonous. Language models could not produce valid plans in Blocksworld, a domain where you stack and unstack labelled blocks under a handful of rules, and they did even worse when the block names and action names were swapped for nonsense words. The best standard model in this paper, LLaMA 3.1 405B, solves 376 of 600 Blocksworld instances zero-shot, and that number was a high water mark.

Then o1-preview arrived on September 12 and Karthik Valmeekam, Kaya Stechly and Subbarao Kambhampati had a result out by September 20. This is the value of a benchmark that predates the model. Nobody designed PlanBench to flatter or embarrass o1. The tests were sitting there, the evaluation code already existed, and the group had a habit of reporting exactly what came out.

What improved

On plain Blocksworld, o1-preview gets 587 of 600 correct zero-shot, which is 97.8 percent. The authors call this a quantum improvement and we think that is fair. o1-mini manages 340 of 600, so the gain is tied to the larger model and the longer reasoning, not to a generic training change. The paper's reading is that o1 has been trained to be an approximate reasoner rather than a retriever, and the Blocksworld numbers alone would support that.

The group is careful about the word approximate. A classical planner, Fast Downward, solves all 600 instances in an average of 0.265 seconds per problem at effectively no cost. o1-preview takes 40.43 seconds on average per Blocksworld instance and, per 100 instances, costs 42.12 dollars against 0.65 for GPT-4o and 0.44 for Claude 3.5 Sonnet. So on the easy version of the domain the new model class is a slower, more expensive, less reliable version of a solver from the last decade.

Where it comes apart

Mystery Blocksworld keeps the same logical structure and renames everything, so a model that has learned the rules should be unaffected and a model that has memorised block-stacking text should fall over. o1-preview drops to 317 of 600, or 52.8 percent, and it gets worse with a worked example in the prompt, 247 of 600 one-shot. On a further randomised version of the obfuscation it falls to 224 of 600. Time per instance roughly doubles to 83 seconds and then to 111.

Plan length is the other axis. On 110 problems that need plans of twenty or more steps, o1-preview scores 23.63 percent. That is the same kind of decay with horizon that the group had documented in ordinary models, just starting from a higher baseline. Whatever o1 is doing during its hidden reasoning, it does not look like search with a guarantee.

The unsolvable instances are the part we keep returning to. Given Blocksworld problems that have no valid plan, o1-preview correctly says so 27 percent of the time, returns an empty plan 19 percent of the time and produces a confident wrong plan 54 percent of the time. On the randomised mystery version it identifies unsolvability 16 percent of the time and hallucinates a plan 79 percent of the time, while also flagging 11.5 percent of solvable problems as impossible. A planner that answers when it should not is a specific failure and a benchmark should have a row for it.

What the design got right

Three features made this evaluation informative in a week. First, the obfuscated variant separates rule following from recall, so a large gain on the surface form can be checked immediately against the underlying skill. Second, the horizon split turns a single accuracy into a curve, and the shape of the curve is what tells you whether the method scales. Third, the unsolvable set catches overconfidence, which accuracy alone hides because a wrong plan and a refusal both score zero.

The cost and latency columns matter too, and the paper puts them beside the accuracy rather than in an appendix. A result like 97.8 percent has a different meaning at 42 dollars per hundred problems and 40 seconds each than it would at the price of a GPT-4o call. Readers of benchmark tables should insist on this.

What we would run next

The obvious follow-up is to give o1 an external verifier and count how many rounds it takes to reach a valid plan on the mystery variant. The authors have done this kind of back-prompting for earlier models and it would tell us whether the reasoning traces are close to correct or simply confident. A second experiment is to sweep the horizon more finely between five and thirty steps and see whether the drop is gradual or a cliff, since that distinguishes a bounded search from a memorised heuristic.

We would also like to see the unsolvable instances broken down by what makes them unsolvable. The current numbers say the model rarely notices, but not why. Until then, the honest summary is the one in the paper's title. Standard models still cannot plan, this one can plan the easy version, and the benchmark that showed both was written before either existed.

Sources

  1. Valmeekam, Stechly and Kambhampati, LLMs Still Can't Plan, Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench (arXiv 2409.13373)