KISS Sorcar Beats Three Coding Harnesses on Terminal-Bench 2.0

Koushik Sen, UC Berkeley

The short version

A group at UC Berkeley and Arena Intelligence recently asked how much the harness around a coding model matters. They held the model fixed, swapped between Claude Code, Codex CLI and Pi, and found that the harness changes what you pay far more than it changes what gets solved. We had a harness of our own, built on the opposite hunch, so we added it to their Terminal-Bench 2.0 table. Averaged over their seven models, it solved 75.6% of attempts. Pi got 70.0%, Codex CLI 65.7% and Claude Code 65.1%. Thirty tasks isn’t many, though, and that pooled gap sits inside the noise. So we reran Pi ourselves on Claude Fable 5, on the same tasks, priced from the same table, with no turn cap for either agent. KISS Sorcar solved 80.0% of attempts to Pi’s 68.9%, an 11-point gap that holds up statistically, for five cents more per attempt. On the 59 Terminal-Bench 2.0 tasks the study didn’t sample, it solved 79.1%.

What HarnessTax found

Every coding agent is two things: a model and the harness wrapped around it. The harness is the loop that calls the model, the tools it exposes, the instructions it sends along, and all the bookkeeping in between. Leaderboards mostly report the pair as one number. Melissa Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica and Matei Zaharia pulled the two apart. Their HarnessTax post takes seven models (Claude Fable 5, Opus 4.8, Sonnet 4.6, Haiku 4.5, GPT-5.6 Sol, GPT-5.6 Luna and Kimi K3), runs each one under three harnesses (Claude Code, Codex CLI and Pi), and scores them on 30 tasks sampled from SWE-bench Lite and 30 from Terminal-Bench 2.0, three attempts per task, with each benchmark’s official verifier doing the grading. Every model-harness pair gets the same 100-turn cap and the same “high” reasoning effort.

Three findings come out of that grid:

  1. The harness moves cost more than it moves correctness. On Terminal-Bench 2.0, switching harness shifts the success rate by about ±5 points on average, while Claude Code costs around 1.5× what Pi does.
  2. A minimal harness holds its own. Pi hands the model four tools (read, write, edit, bash) and not much else, and it lands on the cost-success frontier of both benchmarks.
  3. A model doesn’t necessarily do best in its own vendor’s harness. In nine of the twelve model-benchmark comparisons for the Anthropic and OpenAI models, some other harness had the highest success.

The name comes from the first finding. If you accept a coding agent’s default harness without trying the alternatives, the authors say, you are probably paying a harness tax, meaning more money for the same result.

We have no quarrel with the second and third findings. The first one we wanted to poke at, because our harness was built on a different assumption. We had been betting that a prompt full of solid engineering rules, backed by a little code, changes what the model gets right as well as what it costs, and the study’s published tasks and tables let us test that bet on the same footing.

What we ran

KISS Sorcar is an open-source general-purpose agent and IDE assistant. We have spent the last six months using it to build itself. It consists of five small Python agent classes and a system prompt that is mostly rules about how to work: read a file before you edit it, reproduce a bug before you fix it, verify every requirement from a fresh shell before you say you’re done. The paper about it covers a study of our own task database, an ablation of that prompt, a cross-vendor reviewer experiment and two systems case studies (a key-value store and a TPC-H engine). What it didn’t have was a head-to-head with other harnesses on the same tasks. HarnessTax publishes its task list, its per-model numbers and its protocol, which is what you need for that.

We picked Terminal-Bench 2.0 rather than SWE-bench Lite because its tasks run longer. For the benchmark we stripped the general-purpose material out of the system prompt and kept the coding rules. That left 727 words, down from the 3,851 words of the full SYSTEM.md.

Two of our settings differ from the study’s, and you should keep them in mind when you read the table. We set no turn cap and no wall-clock limit; an attempt ends when the agent calls finish or when it has spent $50. We also sent the providers’ default request parameters, where the study ran at high effort. The tasks, the attempt count and the scoring match the study: the same 30 tasks, three attempts each, graded by the Terminal-Bench 2.0 verifier. A task’s three attempts are averaged first, then the 30 tasks; the pooled figure averages the seven per-model results; the intervals are percentile bootstraps over 10,000 resamples of the tasks. All 630 attempts ran on 25 September 2026, and every one of them is in the table.

How it came out

Table 1 puts the four harnesses next to each other. The three pairs of columns on the right are the study’s published numbers; the pair on the left is ours.

Table 1. Terminal-Bench 2.0, the study’s 30-task sample, three attempts per task: percentage of attempts solved and cost per attempt in USD. KISS Sorcar: 630 attempts, no turn cap, $50 budget. The other three: as published, 100-turn cap, high effort. Bold marks the best point estimate per row.
ModelKISS SorcarPiCodex CLIClaude Code
solved$/att.solved$/att.solved$/att.solved$/att.
Claude Fable 580.01.5071.11.0872.20.9875.61.55
Claude Opus 4.874.41.2172.20.7672.20.8568.90.90
Claude Sonnet 4.676.72.2365.60.6163.30.5562.20.67
Claude Haiku 4.551.10.3947.80.2531.10.2141.10.26
GPT-5.6 Sol85.60.6983.30.4278.90.7671.11.35
GPT-5.6 Luna78.90.0576.70.0472.20.0670.00.10
Kimi K382.21.1473.30.3870.00.4566.70.52
Mean of 7 models75.61.0370.00.5165.70.5565.10.77
0% 20% 40% 60% 80% 100% 80 71 72 76 Fable 5 74 72 72 69 Opus 4.8 77 66 63 62 Sonnet 4.6 51 48 31 41 Haiku 4.5 86 83 79 71 GPT-5.6 Sol 79 77 72 70 GPT-5.6 Luna 82 73 70 67 Kimi K3 76 70 66 65 Mean of 7 KISS Sorcar Pi Codex CLI Claude Code
Figure 1. The success column of Table 1 as bars. KISS Sorcar (blue) has the highest point estimate on each of the seven models. The gap over Pi is largest on Sonnet 4.6 (+11.1) and Kimi K3 (+8.9) and smallest on Opus 4.8 and Luna (+2.2 each) and Sol (+2.3).

Averaged over the seven models, KISS Sorcar solved 75.6% of attempts, with a 95% interval over tasks of 63.8 to 85.9. That is 5.6 points above Pi and 10.5 above Claude Code. On each of the seven models taken individually, it has the highest point estimate of the four. If we treat the solves that took more than 100 turns as failures, to approximate the study’s cap, the pooled rate drops to 74.1%, which is still above the other three.

Thirty tasks is not a lot, and the intervals show it. A per-model interval spans 27 points on average, so no single row of Table 1 settles anything on its own, and the pooled interval of 63.8 to 85.9 contains all three published means. The comparison with the published table also differs from the study in several ways at once: date, price list, turn cap and prompt. On price lists, the study priced its runs on a September 1 list and we priced ours with our framework’s table, which for Kimi K3 mixes two providers, so we don’t draw any cost ratio from this table.

30% 40% 50% 60% 70% 80% 90% $0.05 $0.1 $0.2 $0.5 $1 $2 cost per attempt (USD, log scale) Fable 5 Opus 4.8 Sonnet 4.6 Haiku 4.5 GPT-5.6 Sol GPT-5.6 Luna Kimi K3 KISS Sorcar Pi Codex CLI Claude Code
Figure 2. Table 1 as a cost-success plot on a log cost axis. One diamond per model for KISS Sorcar; the gray lines join it to the same model under Pi, Codex CLI and Claude Code. Every diamond sits above its model’s other three points and, on the published price lists, to the right of Pi.

Rerunning Pi ourselves

The way to remove most of those differences is to rerun the competitor yourself, so we did that for one row. We ran Harbor’s Pi agent (Pi 0.87.1) on Claude Fable 5 on 26 September 2026, the day after our own run, on the same 30 tasks, three attempts each, at the study’s high-effort setting, with no turn cap and no wall-clock limit. Pi has no budget cap, so an attempt ends when the model stops. We priced its token counts with the same table as ours. All 90 attempts count. We killed one after two hours, while Pi was waiting on a training command it had started; that attempt is counted as failed, with its spend.

Pi solved 62 of the 90 attempts, or 68.9% (interval 52.2 to 84.4), at $1.45 and 14.5 turns per attempt. KISS Sorcar on the same model solved 80.0% at $1.50 and 15.3 turns. Both harnesses tried every task, so the comparison is paired; we resample the tasks once and apply the same resample to both harnesses. That gives a success gap of 11.1 points in KISS Sorcar’s favor, with a 95% interval of +3.3 to +20.0. The cost difference is five cents per attempt, with an interval of −$0.71 to +$0.70 that straddles zero.

-5 +0 +5 +10 +15 +20 +25 no difference +11.1 points [+3.3, +20.0] Success gap, KISS Sorcar minus Pi, on the 30 tasks (percentage points; 95% paired bootstrap) KISS Sorcar higher: 7 tie: 23 Pi higher: 0 Per-task tally over the 30 tasks (exact sign test p = 0.016)
Figure 3. The paired comparison on Claude Fable 5. Top: the success gap and its bootstrap interval, which excludes zero. Bottom: KISS Sorcar has the higher task mean on 7 tasks, Pi on none, and 23 are ties, 20 of them solved by both on every attempt.

Pi’s three attempts on feal-linear-cryptanalysis all died with an API error and no answer submitted; two hit a Pi client bug on a mid-output model fallback, and the third hit a content filter. Drop that task and the gap is +8.0 points (6 tasks to 0, p = 0.031). We didn’t rerun Codex CLI or Claude Code, so the paired claim is about Pi, on one model.

The 59 tasks the study left out

A harness that has been tuned on a public sample can look good on that sample and nowhere else. To check, we ran the same agent and prompt on Claude Fable 5 on the 59 Terminal-Bench 2.0 tasks the study didn’t sample, three attempts each, under the same protocol. The prompt was frozen before the run and we didn’t touch it between the two sets.

It solved 79.1% of attempts on the unsampled tasks, and 79.4% across all 89. Those figures count one stopped attempt, on extract-moves-from-video, as failed; it left no result, so we treat it the way we treated Pi’s stopped attempt. If you look only at the 176 attempts that ran to completion, the unsampled tasks come out at 79.7% (interval 70.1 to 88.1) against 80.0% on the study’s 30. The 30-task rate is 0.3 points higher, but the interval on that difference, −15.8 to +15.4, is far too wide to call it a real gap. Over all 89 tasks the completed-attempt rate is 79.8% (71.9 to 86.9).

These tasks cost a little more, $1.60 per attempt against $1.50. The attempts ran longer (17.2 turns against 15.3) but were cheaper per turn ($0.089 against $0.098). The rest of the gap comes from giving one task a full task’s weight: the two attempts of extract-moves-from-video that did finish averaged $13.27, and that task counts as much as any other in the average.

What the trajectories look like

The average attempt took 27.9 turns (the longest, 276) and 17.2 minutes. Twenty-eight attempts went past the study’s 100-turn cap, and 9 of those were solved. A cap costs solves, which is why we gave the 74.1% figure next to the 75.6% above. Three attempts cost more than $15 (the most expensive, $33.64), and two of those were solved.

Of the 630 attempts, 250 received at least one harness note about a still-running process or a changed input file, 901 notes in total. Terminal-Bench tasks routinely start a server, a build or a training job that is still running when the agent reads the next result, and the verifier looks at files the agent may have brushed against by accident. A one-line note puts both facts in front of the model instead of leaving it to guess. The harness refused 15 tool calls, in 15 different attempts. Fourteen were run_agent calls, which have no channel agent to reach inside a container; the other was a shell command that the destructive-command guard misread.

The prompt itself is short, and the full text is archived with the records at papers/kisssorcar/evidence/tb2_prompt.txt. Its six rules, paraphrased:

  1. Understand the task before acting: list what “done” means for every item it names.
  2. Look before you change; never modify the task’s input files, and experiment on copies instead.
  3. Work in small verified steps; reproduce a bug before fixing it; never replace something that works with something unverified.
  4. Verify before you finish, from a fresh shell, requirement by requirement; never weaken a test or threshold to pass a check.
  5. Leave the environment as the task expects to find it.
  6. Be honest at the end: success=True only when every requirement is met and checked.

Cost depends on the number of turns and on what goes out with each one: the prompt, the tool schemas and the notes. The paired run on Claude Fable 5 is the only place we can compare that cost with Pi on one price list, and there the difference was five cents per attempt.

Where this leaves the harness tax

HarnessTax’s central claim is that on these benchmarks the harness moves cost more than success, and for Terminal-Bench 2.0 it puts the harness effect within about ±5 points. Our pooled gap over Pi is 5.6 points, right at the edge of that band. Our paired gap on one model is 11 points, with an interval that doesn’t include zero. The study stands; the band a fourth harness draws is wider than the one three harnesses drew.

The study describes its three harnesses mostly in terms of how much they send the model. Pi sends four tools and little else; on SWE-bench Lite, Claude Code’s first call carries more than ten times Pi’s context. Ours is small by that measure too. What it adds is a short account of how to work, and a note whenever the model can’t see something important in a shell result, such as a process it started that is still running, or a file the verifier is going to inspect that has changed. We have not tested whether those two things are what moved the numbers. The all-model comparison also changes effort, cap and date, and the paired run leaves effort and prompt tangled together.

The HarnessTax authors leave room for this when they close their own post. For day-to-day tasks, they write, a coding agent is essentially an interface to model intelligence, but harder problems “may still benefit from harnesses that provide structured guidance.” Terminal-Bench 2.0 tasks are single-session, but they involve background processes and container state that a model with four tools and no notes has to infer. Our harness supplies 727 words of working rules and a one-line note in the tool result, and on this benchmark the harness that carries them came out ahead. Whether they are the reason is the open question. Settling it means varying the prompt, the notes and the guard one at a time with everything else held fixed, which we have not done.

We can also add a fourth data point to the study’s second and third findings. A simple harness is competitive; ours is a loop, six tools and a prompt. Models also do fine outside their vendor’s harness; every Anthropic and OpenAI model in Table 1 has its best point estimate under a harness that neither vendor wrote.

Things to keep in mind

The paired run matches tasks, scoring and price list and drops the turn cap for both agents, but it still differs from a clean A/B test. Pi ran at high effort with its own prompt and no budget cap, the day after our run, and only Pi was rerun, on one model. The paired interval and the sign test are the results that stand on their own; the pooled interval over seven models covers all three published means.

This benchmark exercises the loop, the tools and the coding rules. The discovery and adversarial-testing procedures, the memory, the reviewer and the IDE features that the rest of the paper is about were switched off, so nothing here speaks to them.

Credit

This comparison was possible because the HarnessTax authors published their task sample, their per-model tables and their protocol.

The full method, the intervals and the rest of the evidence for KISS Sorcar are in KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant, Section 5. Terminal-Bench 2.0 is by Merrill et al., arXiv:2601.11868. Pi is the open-source coding agent at github.com/earendil-works/pi.