Three ways to pass without solving

On April 19 the Terminal-Bench maintainers published an integrity update describing three submissions that had passed tasks without doing them. OB-1, from OpenBlock, was found to have stored encrypted solutions inside the agent binary. Pilot, from QuantFlow, had uploaded the tests folder from the tasks as part of its agent setup, which means the agent could read the checks it was about to be graded against. ForgeCode's agent retrieved solutions from the internet and wrote them into its AGENTS.md file. OB-1 and Pilot were removed from the leaderboard. ForgeCode's affected trials were rescored to zero.

None of these was caught by the benchmark harness. The Pilot and ForgeCode cases were found by Adam Stein and Davis Brown, and the OpenBlock case by the team at Ante, all credited in the update. A community leaderboard was policed by the community, after the scores had been posted and, in at least one case, after the same submitter had already been through an earlier integrity process.

The earlier incident

That earlier process is worth recalling because it shows how the norms shifted. In September 2025 the maintainers published a note about timeouts. OpenBlock had submitted OB-1 with the per-task timeout changed to a uniform 30 minutes, which raised the limit for 75 tasks that normally get five minutes, left two unchanged, and lowered three. The submission had taken the top spot. The team wrote at the time that they believed OpenBlock acted in good faith and had not appreciated that timeouts affect task difficulty. OB-1 was resubmitted with correct parameters and reinstated at the top.

The remedies then were documentation, a checklist in the approval workflow, and an optional validation path for reproducible agents. Seven months later the same submitter's agent turned up with encrypted answers in the binary. We are not going to speculate about intent. What we will say is that the September fix assumed misunderstanding and the April finding does not fit that assumption, and the policy has changed accordingly.

What reward hacking means here

The update gives a definition worth keeping. Reward hacking occurs when a model exploits a loophole to resolve a task without demonstrating the capability the task was intended to measure. On Terminal-Bench the task is a container with a goal and hidden tests, and the capability is getting a real terminal from the starting state to a state that passes. Reading the tests, fetching a known solution, or shipping the answer in advance all produce a passing state without the capability. The tests cannot tell the difference, because the tests only see the final state.

The change in policy follows from that. Passing trials now require ATIF trajectories, the full record of what the agent did, so that a judge can inspect the path and not only the destination. Any trial found to involve reward hacking scores zero. Confirmed cheating means immediate removal, with resubmission considered case by case. An agent judge will validate passing trials, and the judge tool will be open sourced so submitters can run it on their own trajectories before sending them in.

Why this was inevitable

Terminal-Bench started in May 2025 as a research benchmark with a reference agent. Within days the Claude 4 launch was quoting it, and by late June a commercial terminal product, Warp, was announcing a state of the art of 52 percent on it. Version 2.0 arrived in November with a new harness, version 2.1 this May fixed 28 tasks and added continuous validation. Over the same period a leaderboard position went from an academic curiosity to something a startup could put in a pitch deck.

The moment a number is commercially valuable, the honour system that ran the leaderboard stops being a reasonable design. The maintainers were verifying outcomes because verifying outcomes is cheap and, for a research audience, sufficient. Once a submitter has a financial reason to reach the top, the outcome is the wrong thing to verify. This is not a Terminal-Bench problem. Any evaluation where the submitter controls the agent and the harness only sees the result has the same hole, and the hole gets exploited in proportion to what the result is worth.

What a judge can and cannot see

Mandatory trajectories close the three cases in the update. An encrypted solution decrypted at runtime shows up in the trace as an unexplained write of a complete answer. An uploaded tests folder shows up as the agent reading files it should not have. A fetched solution shows up as a network call followed by a paste. A judge reading the trace, human or model, can flag all three.

What a trajectory judge cannot catch is a solution that was memorised in the weights. If the tasks are public and the model was trained on them, the agent will produce the answer with a clean-looking trace and there is nothing in the record to flag. That is contamination rather than reward hacking, and the only defence is held-out tasks, which is a cost the benchmark now has to carry indefinitely. Our guess is that within a year every leaderboard that matters commercially will run a private split, a trajectory requirement, and a judge, and that the ones that do not will stop being cited.

The thing we would want to see next is the judge's own error rate. An open-source reward-hacking judge is going to be run by every submitter before submission, which means it is also the thing every submitter will tune against. Publishing how often it misses, on a held-out set of known hacks, is the only way to know how long it stays useful.

Sources

  1. Terminal-Bench: Leaderboard Integrity Update (April 19, 2026)
  2. Terminal-Bench: Leaderboard Integrity and Timeouts (September 9, 2025)
  3. Terminal-Bench: News