The problem it is solving

Harbor-Index came out on July 7 from Lin Shi, Haowei Lin and colleagues on the Terminal-Bench team, and the motivation is a cost line. The team has adapted 54 benchmarks into their Harbor evaluation framework, and running one agent across all of them costs, by their figures, 226 billion tokens and about 300 thousand dollars in compute. Nobody can afford to do that for every new model and every new scaffold. The index is their attempt to compress that evaluation into something that costs a frontier agent a few hundred dollars per run.

The compression is severe. From 6,627 candidate tasks they kept 82. The interesting question for anyone who builds evaluations is how those 82 were chosen, because the method is the contribution.

Three filters

The first filter is difficulty. Every candidate task was run against three frontier models, Claude Opus 4.6, GPT-5.4 and Gemini 3.1 Pro, in two scaffolds with three repetitions each, for 18 trials per task. Tasks with a pass rate of 34 percent or more across those trials were dropped. That took the pool from 6,627 to 1,311.

The second filter is the one we find most instructive. An LLM auditor read each surviving task and checked two things. Does the verifier actually test what the instruction asks for, and does the difficulty come from the reasoning the task requires rather than from a brittle environment, a flaky dependency or an ambiguous spec. A hard task that is hard because it is broken is worse than an easy one, since it rewards whichever agent happens to be lucky about the breakage. That pass cut the pool to 307.

The third filter is people. Fourteen independent reviewers re-audited the 307 with the same criteria, and a senior panel made the final selection. Then came repeated cycles of running frontier models on the survivors, finding the ways tasks could be gamed or misread, and fixing them. The outcome is 82 tasks drawn from 29 of the original 54 benchmarks, across seven domains. Software engineering contributes 31, scientific research 16, agents and tools 14, knowledge 9, mathematics and reasoning 7, data and analytics 3, and safety and security 2.

What the leaderboard says

Scoring is pass at one across 1,476 rollouts, binary with no partial credit, one rollout per agent and model pair per task. The top of the board is GPT-5.5 running in the Codex CLI at 28.1 percent, then Claude Opus 4.8 in Claude Code at 20.7. Nine agent and model combinations are listed with cost alongside score, and none clears 30 percent. That was the design target, and it means the suite has headroom for a few generations rather than a few months.

The finding that will get argued about is the scaffold comparison. Terminus-2 is the team's own cross-vendor scaffold, restricted to bash so that every model runs in the same environment. Native command line tools from the model vendors beat it. The native agents used about a third fewer tool calls and more than 40 percent fewer output tokens per trial, and they timed out on 26 percent of runs against 42 percent for Terminus-2. The team's explanation is partly that native tools include image reading and web search that a bash-only environment lacks, so some tasks are simply not reachable from Terminus-2.

Curation is the scarce skill

We have written before that benchmark saturation is mostly a story about test sets leaking and tasks being easier than they looked. This release is a demonstration that the fix is labour. Of the 6,627 tasks that 54 benchmark authors published as measuring something, fewer than a fifth were hard for current models, and of those, fewer than a quarter survived an audit for whether the difficulty was real. The audit was done by a model first and by people second, and the people were the expensive step.

That has a consequence for how the field should count evaluation work. Writing 500 tasks is now cheap. Confirming that 80 of them measure what they claim, across three frontier models and two scaffolds with the fixes iterated until the tasks hold, is the part that takes a team and a budget. If Harbor-Index holds up, its value is that a small lab can get a credible number for a few hundred dollars instead of building this pipeline itself.

The scaffold result has a consequence too. If native CLIs outperform a shared scaffold by that margin, then a leaderboard that fixes the scaffold for fairness is measuring a capability that no user experiences, and one that fixes the model is confounded by the vendor's tooling. There is no clean separation. Reporting both, as the index does, is the honest option.

What we would check

Two things worry us and both are testable. First, 82 tasks is small, and pass at one on a binary metric gives wide intervals. We would want to see the variance across repeated rollouts of the same agent before treating a gap of a few points as real. Second, the difficulty filter used three specific models, and a task that is hard for those three may be hard for reasons that a fourth model does not share. Running the filter again with a different trio and checking how much the 82 changes would tell us whether the index measures difficulty or measures those models.

The experiment we would most like to see is longitudinal. Keep the 82 fixed, rerun every new agent as it ships, and publish which tasks fall first. If the ones that fall are concentrated in one domain, the index is telling us something about where progress is. If they fall uniformly, it is telling us the suite was well balanced. Either way we learn more from the order in which a curated set gets solved than from the fact that it eventually is.

Sources

  1. Harbor-Index
  2. Terminal-Bench news